CEREBRAL STRATUM Help

ADR-0010: Cross-Region Observability Egress Boundary

Field

Value

Status

Accepted

Date

2026-09-17

Author

Alex Henshaw

Relates to

cerebralstratum ADR-0001 (Network Topology and Multi-Region Ingress), ADR-0005 (Data Classification and Control Plane Residency); cerebralstratum-backend ADR-0002 (Observability and Fleet Dashboard Embedding)

Tracked in

CSPROD-206 (drafting); implementation Epic to be filed on merge

1. Context

ADR-0005 (Data Classification and Control Plane Residency) named observability egress as a sovereignty risk category but explicitly left it unresolved: "an open cross-region observability egress boundary requiring explicit review of what Grafana Alloy actually ships centrally versus what stays region-local." This ADR is that follow-up.

Platform ADR-0001 (Network Topology and Multi-Region Ingress Strategy) already commits to the shape of the collection mechanism:

Layer

Signal

Collection mechanism

Infrastructure

Node metrics, resource utilisation

Grafana Alloy (OpenShift node exporters)

Platform

PostgreSQL metrics

Grafana Alloy scraping PostgreSQL views (read-only monitoring role, 30s interval)

Application

Traces, metrics, logs

OpenTelemetry SDK (quarkus-opentelemetry), exported via OTLP to Grafana Cloud

Networking / Mesh

Envoy access logs, mesh metrics, request traces

OSSM telemetry pipeline → Grafana Alloy → Grafana Cloud

...and states "Grafana Alloy is deployed as a DaemonSet/Deployment on OpenShift, collecting from both the application OTLP endpoint and the OSSM telemetry pipeline" — but leaves the DaemonSet-vs-Deployment choice, the per-region topology, and (critically, per ADR-0005) the tenant-identifying-content scrub boundary undecided.

This gap is not hypothetical — it is live. A 2026-09-16 investigation (Grafana Cloud showing metrics only from a developer laptop, none from QA or production) traced back to:

  • cerebralstratum-backend's (and device-registrar/notification-dispatcher/device-simulator's) %prod Quarkus profile enables quarkus.otel.traces/metrics/logs but never sets quarkus.otel.exporter.otlp.endpoint. Only %dev sets it (pointed directly at Grafana Cloud's OTLP gateway, authenticated via a header sourced from 1Password through hack_modules.sh — a developer-laptop-only credential path). In %prod, the exporter silently falls back to Quarkus's default http://localhost:4317, which nothing is listening on.

  • Direct cluster inspection (oc against the prod cluster, all namespaces) confirms no Grafana Alloy Deployment, DaemonSet, or Helm release exists anywhere — not in production-cerebral-stratum, not in qa-cerebral-stratum, not as a cluster-wide observability component.

  • Neither the infra repo nor the gitops repo (checked against origin/main in both) contains any Alloy manifest, Helm values, or ArgoCD Application for it.

  • The gitops Rollout specs for cerebralstratum-backend and cerebralstratum-notification-dispatcher, in both QA and production, carry zero QUARKUS_OTEL_* environment variables.

So today's actual state is simpler than "an unresolved sovereignty boundary": QA and production ship no application telemetry to Grafana Cloud at all, because Alloy was never built, and the app-side OTLP endpoint config for a non-dev profile was never finished either. The sovereignty question ADR-0005 raised is real and still needs answering, but it's currently moot in the sense that nothing egresses yet — which is the opportunity to get the boundary right before turning the pipe on, rather than auditing it after the fact.

2. Decision

Regional Alloy, not global; app→Alloy is local-only; Grafana Cloud credentials live only in Alloy; regions that can't reach Grafana Cloud (or shouldn't) get a self-hosted Grafana OSS stack instead.

2.1 Topology: one Alloy Deployment per region

Deploy Grafana Alloy as a namespaced Deployment per region — hub (au-canberra-core, on-premises) plus spokes us-east-2, eu-west-2, ap-southeast-2 today — not a single global instance and not a bare per-node DaemonSet for application/mesh telemetry. Each region's Alloy is reachable only via an in-cluster Service DNS name local to that region/namespace group (e.g. grafana-alloy.<region>-cerebral-stratum.svc.cluster.local:4317).

Applications send OTLP to their own region's Alloy — never across a region boundary. This is set via quarkus.otel.exporter.otlp.endpoint, supplied as an environment variable per-environment in the gitops Rollout manifests (not hardcoded in application.yml, so the same container image is region-portable).

2.2 Credential centralization

Each regional Alloy instance holds the sole export credential for its region — Grafana Cloud remote_write /OTLP credential for the three AWS spokes, or the local Grafana OSS stack's credential for au-canberra-core (see §2.5) — injected via the existing 1Password Operator/Vault pattern already used elsewhere in gitops (prod-core-secrets). Application pods in QA/prod never receive a Grafana Cloud credential — that pattern (QUARKUS_OTEL_EXPORTER_OTLP_HEADERS via hack_modules.sh) stays strictly %dev /developer-laptop-only, as it is today.

2.3 Scrub-then-forward boundary (the actual sovereignty decision)

Before any signal leaves the region, the regional Alloy pipeline classifies it against ADR-0005's two-tier model:

  • Aggregate/non-attributable metrics (JVM stats, HTTP latency histograms, error rates, resource usage — no device_id/tenant_id /request-path-with-identifier content) ship centrally to Grafana Cloud without restriction. This is the ADR-0005 "aggregate/anonymised operational metrics" carve-out.

  • Traces and logs are higher risk — span attributes and log fields can carry device_id, tenant_id, or identifier-bearing request paths (ADR-0005's named risk). An OTel Collector transform/attributes processor stage in each regional Alloy pipeline strips or hashes these fields before central export. Signals that cannot be safely scrubbed (or haven't been reviewed yet — see Open Items) are not exported centrally; they are retained in the region's own Grafana OSS stack (see §2.5) instead of Grafana Cloud.

  • Mesh telemetry (Envoy access logs/metrics via OSSM, per platform ADR-0001) feeds the same regional Alloy instance and is subject to the identical scrub-then-forward rule — access logs carry request paths and must not bypass the boundary just because they originate from the mesh rather than the app.

2.4 PostgreSQL view scraping (backend ADR-0002) is unaffected

The existing cerebralstratum-backend ADR-0002 design (Alloy scraping per-tenant-schema PostgreSQL views, tenant_id promoted to a label) already produces tenant-labelled fleet metrics by design, for customer-facing dashboards accessed via scoped, per-tenant Grafana embed URLs (customer dashboards are inherently single-tenant-scoped at the presentation layer). This ADR does not change that flow. It governs the separate, lower-level application/infra/mesh telemetry path (JVM, HTTP, Envoy, logs) that platform ADR-0001 describes, which has no equivalent per-tenant scoping and is the path actually implicated in the ADR-0005 open item.

2.5 Regional storage backend, and the au-canberra-core exception

The three AWS spoke regions (us-east-2, eu-west-2, ap-southeast-2) run on ROSA HCP, so any scrub-failed traces/logs that need a region-local sink land on Loki/Tempo backed by native S3 — no additional storage backend is required, since S3 is already there. Loki and Tempo themselves are still net-new services to deploy and operate per spoke region, though: confirmed via a check of both infra and gitops that neither repo references Loki, Tempo, or MinIO anywhere today. That deployment work is called out explicitly in Forward Pointers below rather than folded into the Core-specific bullet, so it doesn't get silently dropped when this becomes implementation tickets.

au-canberra-core is different in three ways that this ADR treats as a single, related decision:

  • It's on-premises SNO with no native S3. Red Hat OpenShift Data Foundation (Ceph) was considered and rejected for this role: on a single-node cluster already running the full app stack, Ceph's resource footprint competes directly with the workloads it's meant to support, and it's a paid OCP add-on rather than bundled entitlement. Object storage is instead provided by MinIO (Apache-licensed, no additional licensing cost, consistent with the platform's existing AGPL/Apache/MIT licensing posture) backing a self-hosted Grafana OSS + Loki + Tempo stack. Given Core's node count, MinIO runs as a single instance, PVC-backed, without erasure coding — this is an accepted non-HA component, consistent with the single-point-of-failure profile Core's SNO topology already carries for the app stack; this ADR does not introduce that trade-off, it inherits it. MinIO's object storage sits on the SSD tier only (Loki/Tempo indexing and recent-data queries are latency-sensitive; the HDD tier is not used unless a future ILM/tiering rule is added for genuinely cold, archival data). Retention on this store is tied to the same purge cadence as the platform's four-state data lifecycle grace period, not an independent retention clock — a concrete TTL is still an open item (see §5).

  • It hosts a full QA and a full production cerebral-stratum stack simultaneously, on top of its hub duties (spoke lifecycle management, image builds via Tekton, ArgoCD hub management). Each app stack gets its own Alloy Deployment in its own namespace (qa-cerebral-stratum, production-cerebral-stratum) — QA and prod telemetry must not share a pipeline or credential on Core, exactly as they don't share one across the AWS regions.

  • Core's production stack is treated as carrying real tenant-classified data, even though its actual traffic is BlueGuardian-employee-only. This is deliberate: it means the scrub-then-forward mechanism gets exercised end-to-end against genuine tenant-schema data before any customer-facing spoke goes live, rather than being validated for the first time against real customer traffic.

Reachability. Core's self-hosted Grafana OSS is exposed using the platform's existing pattern for Core-hosted services — Cloudflare Tunnel + a published application, routed through the OSSM ingress gateway into the mesh — rather than a new exposure mechanism. No bespoke ingress path is introduced for observability.

3. Alternatives Considered

Alternative

Reason rejected

Every pod ships OTLP directly to Grafana Cloud with its own credential (today's %dev-only pattern, naively extended to %prod)

Distributes a real Grafana Cloud API credential into every pod across every region; no scrub point for tenant-identifying span/log content; directly the risk ADR-0005 exists to close.

Single global (hub-only) Alloy instance

Every spoke region's app would have to send raw OTLP across the region boundary to reach the hub before any scrub step — crossing the exact boundary ADR-0005 says must not happen for anything carrying tenant-identifying content pre-scrub.

Bare per-node DaemonSet as the sole Alloy topology (literal reading of ADR-0001's "DaemonSet/Deployment")

Scatters the scrub pipeline and the Grafana Cloud credential across every node in a region instead of one namespaced Deployment — larger credential blast radius for no isolation benefit for app/mesh telemetry. (Node-level infra metrics like cAdvisor/node-exporter may still legitimately want DaemonSet placement later — see Open Items — but that's a separate, lower-risk signal class with no tenant-identifying surface.)

ODF/Ceph as au-canberra-core's object storage backend

Resource footprint too heavy for a single-node cluster already running dual app stacks plus hub duties; requires a paid OCP entitlement add-on rather than being bundled. MinIO achieves the same S3-compatible interface at a fraction of the resource and licensing cost.

4. Consequences

Positive

  • Closes the ADR-0005 open item with a concrete, reviewable boundary instead of an implicit assumption.

  • Directly explains and fixes the empty-Grafana-metrics symptom that triggered this investigation.

  • Centralizes the export credential to N regional Alloy instances instead of N×pods.

  • Gives backend/notification-dispatcher/device-registrar/device-simulator a single, consistent pattern for wiring %prod/%qa OTLP export (currently missing in all four).

  • The au-canberra-core self-hosted Grafana OSS stack, built to satisfy the sovereignty boundary, doubles as the entire observability answer for ENTITLEMENT_MODE=community self-hosted operators — no separate design needed for that case.

Negative / Trade-offs

  • Alloy becomes a new per-region component that itself needs monitoring — an Alloy outage in a region is now a telemetry blind spot for that region, with no automatic failover to another region's Alloy (crossing regions to fail over would itself violate the boundary this ADR sets).

  • The scrub pipeline (OTTL/River transform config) becomes sovereignty-relevant code, not just infra config — it needs the same review rigor as application code handling tenant data, not a one-time setup.

  • One additional network hop (app → regional Alloy → Grafana Cloud/regional Grafana OSS) versus direct export, though this is the same shape ADR-0001 already committed to for mesh telemetry.

  • au-canberra-core now runs a materially heavier observability footprint (Alloy ×2, Grafana OSS, Loki, Tempo, MinIO) alongside dual app stacks, image builds, and hub-management duties, all on a single node currently provisioned at 80 vCPU / ~64GB (128GB planned) — see Open Items: this lands on an already-tight memory budget, not spare headroom.

5. Open Items

  • Concrete scrub rules. This ADR sets the policy (aggregate ships, identifying content doesn't without scrubbing) but not the field-by-field OTTL/River transform rules. Needs the concrete trace-span-attribute audit that ADR-0005 also lists as its own open item — what does backend/notification-dispatcher currently attach to spans and logs today, exactly.

  • Node-level DaemonSet exporters (cAdvisor, node-exporter) for pure infra metrics with no tenant-identifying surface — plausibly fine to ship centrally with less scrutiny than app/mesh telemetry, but not yet explicitly scoped in or out of this ADR.

  • gitops placement. Where the regional Alloy Deployment/Service/ConfigMap manifests live in the gitops repo (new observability/ workload dir under each environments/<region>/workloads/?) and which ArgoCD ApplicationSet wires it in — needs an infra PR, not yet drafted.

  • Application-side follow-through. Once the regional Alloy Service DNS name is fixed, cerebralstratum-backend, device-registrar, notification-dispatcher, and device-simulator all need %prod (and a currently-nonexistent %qa) profile blocks updated with quarkus.otel.exporter.otlp.endpoint, and the gitops Rollout specs need the corresponding QUARKUS_OTEL_EXPORTER_OTLP_ENDPOINT env var per region. Tracked as implementation work now that this ADR is accepted, not fixed ad hoc ahead of it.

  • au-canberra-core resource sizing — higher priority than "not yet specced" implies. No explicit resource requests/limits are yet defined for the Core observability stack (2× Alloy, Grafana OSS, Loki, Tempo, MinIO), and direct inspection of node1 shows this isn't landing on spare capacity: cluster-wide memory limits already total ~113% of allocatable (67,640Mi committed vs ~63,740Mi capacity/~59,545Mi allocatable) before any of this stack is added. CPU has real headroom by comparison (~60% of 79,500m allocatable committed). Node capacity reports as 80 CPU via Kubernetes — confirm whether that's 40 physical cores exposed as 80 vCPUs via SMT, since sizing work has to budget against the 80 Kubernetes actually schedules on, not a physical-core count. This needs a memory-limit audit/right-sizing pass on existing workloads and/or the planned 128GB upgrade landing before Core's observability components ship — not deferred as routine follow-up sizing work. No existing YouTrack ticket tracks this; one should be filed alongside this ADR's implementation Epic.

  • MinIO retention TTL. Retention for scrub-failed signals on au-canberra-core's MinIO-backed store is policy-tied to the four-state data lifecycle purge cadence in principle, but a concrete TTL value hasn't been set.

6. Forward Pointers

  • cerebralstratum-backend (and the other three modules') application.yml — %prod /new %qa otel exporter config.

  • gitops repo — Rollout env vars (QUARKUS_OTEL_EXPORTER_OTLP_ENDPOINT) for backend/notification-dispatcher/device-registrar in both environments/qa and environments/production, across all four regions including au-canberra-core.

  • infra repo — new Grafana Alloy Deployment/Service manifests per region, plus 1Password/Vault secret wiring for the export credential (replacing the developer-laptop-only hack_modules.sh/op pattern as the production credential path).

  • infra repo — Loki/Tempo deployment for the three AWS spoke regions (us-east-2, eu-west-2, ap-southeast-2), backed by native S3. Net-new work: confirmed no Loki/Tempo/MinIO reference exists anywhere in infra or gitops today.

  • infra repo — au-canberra-core-specific manifests for Grafana OSS, Loki, Tempo, and MinIO, plus the Cloudflare Tunnel + published application config for reaching Core's Grafana OSS through the OSSM gateway. Gated on the Core resource-sizing open item above landing first.

  • ADR-0005's own open item: formal trace-span-attribute content audit — a prerequisite for writing the scrub rules referenced in Open Items above.

Last modified: 17 September 2026