ADR-0010: Cross-Region Observability Egress Boundary
Field | Value |
|---|---|
Status | Accepted |
Date | 2026-09-17 |
Author | Alex Henshaw |
Relates to |
|
Tracked in | CSPROD-206 (drafting); implementation Epic to be filed on merge |
1. Context
ADR-0005 (Data Classification and Control Plane Residency) named observability egress as a sovereignty risk category but explicitly left it unresolved: "an open cross-region observability egress boundary requiring explicit review of what Grafana Alloy actually ships centrally versus what stays region-local." This ADR is that follow-up.
Platform ADR-0001 (Network Topology and Multi-Region Ingress Strategy) already commits to the shape of the collection mechanism:
Layer | Signal | Collection mechanism |
|---|---|---|
Infrastructure | Node metrics, resource utilisation | Grafana Alloy (OpenShift node exporters) |
Platform | PostgreSQL metrics | Grafana Alloy scraping PostgreSQL views (read-only monitoring role, 30s interval) |
Application | Traces, metrics, logs | OpenTelemetry SDK ( |
Networking / Mesh | Envoy access logs, mesh metrics, request traces | OSSM telemetry pipeline → Grafana Alloy → Grafana Cloud |
...and states "Grafana Alloy is deployed as a DaemonSet/Deployment on OpenShift, collecting from both the application OTLP endpoint and the OSSM telemetry pipeline" — but leaves the DaemonSet-vs-Deployment choice, the per-region topology, and (critically, per ADR-0005) the tenant-identifying-content scrub boundary undecided.
This gap is not hypothetical — it is live. A 2026-09-16 investigation (Grafana Cloud showing metrics only from a developer laptop, none from QA or production) traced back to:
cerebralstratum-backend's (anddevice-registrar/notification-dispatcher/device-simulator's)%prodQuarkus profile enablesquarkus.otel.traces/metrics/logsbut never setsquarkus.otel.exporter.otlp.endpoint. Only%devsets it (pointed directly at Grafana Cloud's OTLP gateway, authenticated via a header sourced from 1Password throughhack_modules.sh— a developer-laptop-only credential path). In%prod, the exporter silently falls back to Quarkus's defaulthttp://localhost:4317, which nothing is listening on.Direct cluster inspection (
ocagainst the prod cluster, all namespaces) confirms no Grafana Alloy Deployment, DaemonSet, or Helm release exists anywhere — not inproduction-cerebral-stratum, not inqa-cerebral-stratum, not as a cluster-wide observability component.Neither the
infrarepo nor thegitopsrepo (checked againstorigin/mainin both) contains any Alloy manifest, Helm values, or ArgoCD Application for it.The
gitopsRollout specs forcerebralstratum-backendandcerebralstratum-notification-dispatcher, in both QA and production, carry zeroQUARKUS_OTEL_*environment variables.
So today's actual state is simpler than "an unresolved sovereignty boundary": QA and production ship no application telemetry to Grafana Cloud at all, because Alloy was never built, and the app-side OTLP endpoint config for a non-dev profile was never finished either. The sovereignty question ADR-0005 raised is real and still needs answering, but it's currently moot in the sense that nothing egresses yet — which is the opportunity to get the boundary right before turning the pipe on, rather than auditing it after the fact.
2. Decision
Regional Alloy, not global; app→Alloy is local-only; Grafana Cloud credentials live only in Alloy; regions that can't reach Grafana Cloud (or shouldn't) get a self-hosted Grafana OSS stack instead.
2.1 Topology: one Alloy Deployment per region
Deploy Grafana Alloy as a namespaced Deployment per region — hub (au-canberra-core, on-premises) plus spokes us-east-2, eu-west-2, ap-southeast-2 today — not a single global instance and not a bare per-node DaemonSet for application/mesh telemetry. Each region's Alloy is reachable only via an in-cluster Service DNS name local to that region/namespace group (e.g. grafana-alloy.<region>-cerebral-stratum.svc.cluster.local:4317).
Applications send OTLP to their own region's Alloy — never across a region boundary. This is set via quarkus.otel.exporter.otlp.endpoint, supplied as an environment variable per-environment in the gitops Rollout manifests (not hardcoded in application.yml, so the same container image is region-portable).
2.2 Credential centralization
Each regional Alloy instance holds the sole export credential for its region — Grafana Cloud remote_write /OTLP credential for the three AWS spokes, or the local Grafana OSS stack's credential for au-canberra-core (see §2.5) — injected via the existing 1Password Operator/Vault pattern already used elsewhere in gitops (prod-core-secrets). Application pods in QA/prod never receive a Grafana Cloud credential — that pattern (QUARKUS_OTEL_EXPORTER_OTLP_HEADERS via hack_modules.sh) stays strictly %dev /developer-laptop-only, as it is today.
2.3 Scrub-then-forward boundary (the actual sovereignty decision)
Before any signal leaves the region, the regional Alloy pipeline classifies it against ADR-0005's two-tier model:
Aggregate/non-attributable metrics (JVM stats, HTTP latency histograms, error rates, resource usage — no
device_id/tenant_id/request-path-with-identifier content) ship centrally to Grafana Cloud without restriction. This is the ADR-0005 "aggregate/anonymised operational metrics" carve-out.Traces and logs are higher risk — span attributes and log fields can carry
device_id,tenant_id, or identifier-bearing request paths (ADR-0005's named risk). An OTel Collectortransform/attributesprocessor stage in each regional Alloy pipeline strips or hashes these fields before central export. Signals that cannot be safely scrubbed (or haven't been reviewed yet — see Open Items) are not exported centrally; they are retained in the region's own Grafana OSS stack (see §2.5) instead of Grafana Cloud.Mesh telemetry (Envoy access logs/metrics via OSSM, per platform ADR-0001) feeds the same regional Alloy instance and is subject to the identical scrub-then-forward rule — access logs carry request paths and must not bypass the boundary just because they originate from the mesh rather than the app.
2.4 PostgreSQL view scraping (backend ADR-0002) is unaffected
The existing cerebralstratum-backend ADR-0002 design (Alloy scraping per-tenant-schema PostgreSQL views, tenant_id promoted to a label) already produces tenant-labelled fleet metrics by design, for customer-facing dashboards accessed via scoped, per-tenant Grafana embed URLs (customer dashboards are inherently single-tenant-scoped at the presentation layer). This ADR does not change that flow. It governs the separate, lower-level application/infra/mesh telemetry path (JVM, HTTP, Envoy, logs) that platform ADR-0001 describes, which has no equivalent per-tenant scoping and is the path actually implicated in the ADR-0005 open item.
2.5 Regional storage backend, and the au-canberra-core exception
The three AWS spoke regions (us-east-2, eu-west-2, ap-southeast-2) run on ROSA HCP, so any scrub-failed traces/logs that need a region-local sink land on Loki/Tempo backed by native S3 — no additional storage backend is required, since S3 is already there. Loki and Tempo themselves are still net-new services to deploy and operate per spoke region, though: confirmed via a check of both infra and gitops that neither repo references Loki, Tempo, or MinIO anywhere today. That deployment work is called out explicitly in Forward Pointers below rather than folded into the Core-specific bullet, so it doesn't get silently dropped when this becomes implementation tickets.
au-canberra-core is different in three ways that this ADR treats as a single, related decision:
It's on-premises SNO with no native S3. Red Hat OpenShift Data Foundation (Ceph) was considered and rejected for this role: on a single-node cluster already running the full app stack, Ceph's resource footprint competes directly with the workloads it's meant to support, and it's a paid OCP add-on rather than bundled entitlement. Object storage is instead provided by MinIO (Apache-licensed, no additional licensing cost, consistent with the platform's existing AGPL/Apache/MIT licensing posture) backing a self-hosted Grafana OSS + Loki + Tempo stack. Given Core's node count, MinIO runs as a single instance, PVC-backed, without erasure coding — this is an accepted non-HA component, consistent with the single-point-of-failure profile Core's SNO topology already carries for the app stack; this ADR does not introduce that trade-off, it inherits it. MinIO's object storage sits on the SSD tier only (Loki/Tempo indexing and recent-data queries are latency-sensitive; the HDD tier is not used unless a future ILM/tiering rule is added for genuinely cold, archival data). Retention on this store is tied to the same purge cadence as the platform's four-state data lifecycle grace period, not an independent retention clock — a concrete TTL is still an open item (see §5).
It hosts a full QA and a full production cerebral-stratum stack simultaneously, on top of its hub duties (spoke lifecycle management, image builds via Tekton, ArgoCD hub management). Each app stack gets its own Alloy Deployment in its own namespace (
qa-cerebral-stratum,production-cerebral-stratum) — QA and prod telemetry must not share a pipeline or credential on Core, exactly as they don't share one across the AWS regions.Core's production stack is treated as carrying real tenant-classified data, even though its actual traffic is BlueGuardian-employee-only. This is deliberate: it means the scrub-then-forward mechanism gets exercised end-to-end against genuine tenant-schema data before any customer-facing spoke goes live, rather than being validated for the first time against real customer traffic.
Reachability. Core's self-hosted Grafana OSS is exposed using the platform's existing pattern for Core-hosted services — Cloudflare Tunnel + a published application, routed through the OSSM ingress gateway into the mesh — rather than a new exposure mechanism. No bespoke ingress path is introduced for observability.
3. Alternatives Considered
Alternative | Reason rejected |
|---|---|
Every pod ships OTLP directly to Grafana Cloud with its own credential (today's | Distributes a real Grafana Cloud API credential into every pod across every region; no scrub point for tenant-identifying span/log content; directly the risk ADR-0005 exists to close. |
Single global (hub-only) Alloy instance | Every spoke region's app would have to send raw OTLP across the region boundary to reach the hub before any scrub step — crossing the exact boundary ADR-0005 says must not happen for anything carrying tenant-identifying content pre-scrub. |
Bare per-node DaemonSet as the sole Alloy topology (literal reading of ADR-0001's "DaemonSet/Deployment") | Scatters the scrub pipeline and the Grafana Cloud credential across every node in a region instead of one namespaced Deployment — larger credential blast radius for no isolation benefit for app/mesh telemetry. (Node-level infra metrics like cAdvisor/node-exporter may still legitimately want DaemonSet placement later — see Open Items — but that's a separate, lower-risk signal class with no tenant-identifying surface.) |
ODF/Ceph as | Resource footprint too heavy for a single-node cluster already running dual app stacks plus hub duties; requires a paid OCP entitlement add-on rather than being bundled. MinIO achieves the same S3-compatible interface at a fraction of the resource and licensing cost. |
4. Consequences
Positive
Closes the ADR-0005 open item with a concrete, reviewable boundary instead of an implicit assumption.
Directly explains and fixes the empty-Grafana-metrics symptom that triggered this investigation.
Centralizes the export credential to N regional Alloy instances instead of N×pods.
Gives backend/notification-dispatcher/device-registrar/device-simulator a single, consistent pattern for wiring
%prod/%qaOTLP export (currently missing in all four).The
au-canberra-coreself-hosted Grafana OSS stack, built to satisfy the sovereignty boundary, doubles as the entire observability answer forENTITLEMENT_MODE=communityself-hosted operators — no separate design needed for that case.
Negative / Trade-offs
Alloy becomes a new per-region component that itself needs monitoring — an Alloy outage in a region is now a telemetry blind spot for that region, with no automatic failover to another region's Alloy (crossing regions to fail over would itself violate the boundary this ADR sets).
The scrub pipeline (OTTL/River transform config) becomes sovereignty-relevant code, not just infra config — it needs the same review rigor as application code handling tenant data, not a one-time setup.
One additional network hop (app → regional Alloy → Grafana Cloud/regional Grafana OSS) versus direct export, though this is the same shape ADR-0001 already committed to for mesh telemetry.
au-canberra-corenow runs a materially heavier observability footprint (Alloy ×2, Grafana OSS, Loki, Tempo, MinIO) alongside dual app stacks, image builds, and hub-management duties, all on a single node currently provisioned at 80 vCPU / ~64GB (128GB planned) — see Open Items: this lands on an already-tight memory budget, not spare headroom.
5. Open Items
Concrete scrub rules. This ADR sets the policy (aggregate ships, identifying content doesn't without scrubbing) but not the field-by-field OTTL/River transform rules. Needs the concrete trace-span-attribute audit that ADR-0005 also lists as its own open item — what does
backend/notification-dispatchercurrently attach to spans and logs today, exactly.Node-level DaemonSet exporters (cAdvisor, node-exporter) for pure infra metrics with no tenant-identifying surface — plausibly fine to ship centrally with less scrutiny than app/mesh telemetry, but not yet explicitly scoped in or out of this ADR.
gitops placement. Where the regional Alloy Deployment/Service/ConfigMap manifests live in the
gitopsrepo (newobservability/workload dir under eachenvironments/<region>/workloads/?) and which ArgoCD ApplicationSet wires it in — needs an infra PR, not yet drafted.Application-side follow-through. Once the regional Alloy Service DNS name is fixed,
cerebralstratum-backend,device-registrar,notification-dispatcher, anddevice-simulatorall need%prod(and a currently-nonexistent%qa) profile blocks updated withquarkus.otel.exporter.otlp.endpoint, and the gitops Rollout specs need the correspondingQUARKUS_OTEL_EXPORTER_OTLP_ENDPOINTenv var per region. Tracked as implementation work now that this ADR is accepted, not fixed ad hoc ahead of it.au-canberra-coreresource sizing — higher priority than "not yet specced" implies. No explicit resource requests/limits are yet defined for the Core observability stack (2× Alloy, Grafana OSS, Loki, Tempo, MinIO), and direct inspection ofnode1shows this isn't landing on spare capacity: cluster-wide memory limits already total ~113% of allocatable (67,640Mi committed vs ~63,740Mi capacity/~59,545Mi allocatable) before any of this stack is added. CPU has real headroom by comparison (~60% of 79,500m allocatable committed). Node capacity reports as 80 CPU via Kubernetes — confirm whether that's 40 physical cores exposed as 80 vCPUs via SMT, since sizing work has to budget against the 80 Kubernetes actually schedules on, not a physical-core count. This needs a memory-limit audit/right-sizing pass on existing workloads and/or the planned 128GB upgrade landing before Core's observability components ship — not deferred as routine follow-up sizing work. No existing YouTrack ticket tracks this; one should be filed alongside this ADR's implementation Epic.MinIO retention TTL. Retention for scrub-failed signals on
au-canberra-core's MinIO-backed store is policy-tied to the four-state data lifecycle purge cadence in principle, but a concrete TTL value hasn't been set.
6. Forward Pointers
cerebralstratum-backend(and the other three modules')application.yml—%prod/new%qaotel exporter config.gitopsrepo — Rollout env vars (QUARKUS_OTEL_EXPORTER_OTLP_ENDPOINT) for backend/notification-dispatcher/device-registrar in bothenvironments/qaandenvironments/production, across all four regions includingau-canberra-core.infrarepo — new Grafana Alloy Deployment/Service manifests per region, plus 1Password/Vault secret wiring for the export credential (replacing the developer-laptop-onlyhack_modules.sh/oppattern as the production credential path).infrarepo — Loki/Tempo deployment for the three AWS spoke regions (us-east-2,eu-west-2,ap-southeast-2), backed by native S3. Net-new work: confirmed no Loki/Tempo/MinIO reference exists anywhere ininfraorgitopstoday.infrarepo —au-canberra-core-specific manifests for Grafana OSS, Loki, Tempo, and MinIO, plus the Cloudflare Tunnel + published application config for reaching Core's Grafana OSS through the OSSM gateway. Gated on the Core resource-sizing open item above landing first.ADR-0005's own open item: formal trace-span-attribute content audit — a prerequisite for writing the scrub rules referenced in Open Items above.