CEREBRAL STRATUM Help

ADR-0007: Consent-Gated Device Telemetry Mirroring from Production to QA

Field

Value

Status

Proposed

Date

2026-09-12

Author

Alex Henshaw

Relates to

cerebralstratum ADR-0001 (Network Topology and Multi-Region Ingress), ADR-0005 (Data Classification and Control Plane Residency), ADR-0006 (Keycloak Deployment Topology); cerebralstratum-backend ADR-0005 (UMA 2.0 Device Resources)

Tracked in

(Epic to be filed on merge)

Terminology: QA workload environment vs. QA cluster

This ADR concerns the QA workload environment — a namespace on the same physical cluster as production, used for feature and telemetry-mirroring testing. This is distinct from the QA clusters (the on-prem SNO ×3, one per region), which are separate physical hardware used for testing infrastructure config changes (Operator upgrades, cluster-wide policy, network topology) rather than workload-level feature testing. The regional-boundary enforcement discussed below concerns cross-region cluster boundaries, not the QA-workload/prod-workload distinction within a single cluster.

Context

QA needs realistic device telemetry to validate changes against real-world traffic patterns — volume, message shapes, edge cases — that synthetic or dev-services-generated data doesn't reproduce. The obvious mechanism, Kafka MirrorMaker 2, replicates topics wholesale with no content-level filtering; pointed at production's hono.telemetry.cerebral-stratum topic, it would indiscriminately mirror every customer's device telemetry into QA regardless of consent.

That's a hard blocker on its own terms — customers have not agreed to their location/telemetry data being processed anywhere outside production — and it's also in direct tension with ADR-0005's existing classification: device telemetry is named explicitly as regionally-bound customer data that "does not cross region boundaries in any form, including transiently." Any design here has to satisfy customer privacy as a first-class constraint, not something bolted on after the fact.

Two categories of device legitimately need to reach QA:

  1. BlueGuardian's own prototype/test-fleet devices — internal, no customer consent question, developers should be able to add/remove them freely.

  2. Customer devices, opted in — a customer voluntarily agrees to let their device's telemetry help validate the platform in QA. This must be genuinely opt-in (deny-by-default) and revocable.

Decision

A dedicated bridge service, not MirrorMaker 2. A small consumer/producer service — the same architectural shape as the platform's existing Notification Dispatch or Anomaly Detection services, not a generic replication tool — sits between production's and QA's Kafka clusters. It consumes production's hono.telemetry.cerebral-stratum, filters record-by-record against a live consent cache, and republishes only matching records into QA's hono.telemetry.cerebral-stratum.

One attribute, two write paths. Both device categories collapse into a single boolean on the Device entity: telemetryQaMirroringEnabled. The bridge doesn't need to know why a device is eligible, only whether it currently is. Two authorization paths write the same field:

  • A platform-operator-facing endpoint, for BlueGuardian's own test-fleet devices.

  • A customer-facing self-service toggle (device/account settings), for opt-in.

Consent is event-sourced, not queried. Per this platform's own rule that no service bypasses the message bus for cross-service state, the bridge does not query backend's Postgres directly. Backend publishes every change to telemetryQaMirroringEnabled onto a compacted Kafka topic (e.g. device.qa-mirroring-consent, keyed by device ID; a tombstone represents "not enabled," so compaction naturally prunes revoked entries and the cache defaults to deny for any device with no record at all). The bridge consumes this into a local materialized cache (a Kafka Streams GlobalKTable or equivalent) and filters the telemetry stream against it in real time — no polling, and revocation takes effect on the next message.

Revocation purges, not just stops. Turning telemetryQaMirroringEnabled off must do more than stop future mirroring — already-mirrored telemetry for that device must be purged from QA, consistent with ADR-0005's stance that consent-driven data lifecycle is a real deletion obligation, not a best-effort stop. The consent-revocation event is the trigger for a QA-scoped purge, mirroring the "trigger centrally, execute locally" shape ADR-0005 already establishes for the four-state data lifecycle.

Regional boundary is inherited, not re-litigated — and enforced in code, not just documented. ADR-0005's classification of device telemetry as regionally-bound, never-crosses-a-region data still applies here in full. This decision does not create an exception to it: a region's QA instance may only ever mirror from that same region's production instance. Today this is moot in practice — QA-workload and production-workload both run in the same physical core cluster, with no live multi-region production deployment yet — but the enforcement mechanism must exist now, structurally, rather than being deferred as a documented promise:

  • DNS-zone isolation as the primary guard. The bridge's source and destination Kafka bootstrap addresses are configured exclusively as cluster-internal Service DNS names (<service>.<namespace>.svc.cluster.local), never an external route, load balancer hostname, or IP. Cluster-internal DNS zones don't span physical clusters — two separate ROSA HCP clusters have no shared DNS resolution path between them by default. This means a misconfiguration that points the bridge at another region's Kafka doesn't silently succeed; it fails to resolve. The bridge validates this suffix pattern on startup and refuses to start otherwise.

  • Admission control as a second layer. Once Kyverno and/or RHACS admission control (currently under evaluation) is adopted, a policy rejects any bridge deployment manifest whose Kafka bootstrap configuration references a non-internal address, closing the gap where someone edits the YAML directly rather than going through validated config.

  • Network policy as defense in depth. With OSSM/Istio STRICT mTLS already mesh-wide, a namespace-scoped AuthorizationPolicy restricts the bridge's egress to only the specific KafkaUser /topic ACLs for its own region-pair. Even a compromised or misconfigured bridge pod cannot reach a Kafka cluster it has no legitimate reason to talk to — this is a network-enforcement backstop independent of the DNS-naming argument.

This three-layer enforcement is necessary, not decorative, precisely because the DNS-isolation guarantee is contingent. The moment ADR-0001's multi-region rollout introduces any legitimate cross-cluster network path (e.g. Cloudflare tunnels for control-plane routing), cluster-internal DNS names may become technically reachable across regions for other purposes — a hostname-suffix check alone would no longer be sufficient, since the suffix doesn't tell you which cluster resolved it. At that point the bridge's validation must check the resolved address's cluster identity (or region-scoped service mesh trust domain) rather than the DNS suffix alone, and the admission-control and network-policy layers become the primary guarantees rather than backstops. This dependency is flagged explicitly here so it is not silently missed when ADR-0001 lands — see Forward Pointers.

Alternatives Considered

  • MirrorMaker 2, wholesale. Rejected outright — no content-level filtering; would mirror every customer's data regardless of consent.

  • MirrorMaker 2 + a custom Kafka Connect SMT/predicate for filtering. Considered. Would work, but requires compiling and deploying a custom Java Connect plugin for what is fundamentally simple per-record filtering logic, and doesn't cleanly express the two-actor (developer vs. customer) consent semantics. More operational surface for less flexibility than a purpose-built service.

  • ksqlDB-based filtering. Considered. Introduces an entirely new stateful platform component (a ksqlDB server) solely for this one use case, where a small Kafka Streams/Quarkus consumer-producer achieves the same result using a component shape the platform already runs elsewhere.

  • Bridge queries backend's database directly for consent state. Rejected — violates the platform's stated "no service bypasses the message bus for cross-service state" principle, and would additionally require granting a QA-side (lower-trust) component read access into production's database.

  • Static allowlist only, no customer opt-in. Rejected per explicit product direction — customers should be able to voluntarily help improve the platform, not just have BlueGuardian's own devices eligible.

  • Documented regional-boundary constraint with no code-level enforcement. Considered and rejected — relying on the ADR text alone to prevent a future cross-region mirroring misconfiguration is exactly the kind of latent obligation that gets silently violated once multi-region rollout happens. Enforced via DNS-zone isolation, admission control, and namespace-scoped network policy instead (see Decision).

Consequences

Positive

  • Customer telemetry only ever reaches QA with explicit, revocable, per-device consent — no wholesale or implicit data movement, satisfying the privacy requirement that motivated this decision.

  • The same mechanism serves both BlueGuardian-internal test-fleet devices and genuine customer opt-in, minimizing the design surface (one attribute, one topic, one filter) rather than maintaining two parallel systems.

  • Consent state changes are event-sourced — the compacted topic is itself an audit trail of who had mirroring enabled and when, which matters directly for a customer-privacy-motivated feature.

  • Consistent with the platform's existing service-decomposition pattern (a small, single-purpose Kafka consumer/producer) and its message-bus-only cross-service state rule — no new architectural category introduced.

  • The regional boundary is a structural guarantee (DNS isolation + admission control + network policy) rather than a documented promise, closing the gap between "we said we wouldn't" and "the system can't."

Negative / Trade-offs

  • Mirrored telemetry for an opted-in customer device needs a corresponding device/tenant record in QA for backend to process it meaningfully (schema-per-tenant routing, device metadata lookups) — this ADR does not fully resolve how that metadata reaches QA. Until resolved, mirroring is genuinely useful mainly for BlueGuardian's own test-fleet devices (already dual-provisioned by definition); real customer opt-in devices need this gap closed first. See Open Items.

  • Revocation-triggers-purge is additional implementation surface beyond the mirroring path itself — a QA-scoped purge job triggered off the same consent event, not yet designed in detail.

  • A new always-on service is another component to build, deploy, monitor, and secure (its own KafkaUser /ACLs in both environments, read on two production topics, write on one QA topic).

  • The DNS-suffix check is a point-in-time guarantee, contingent on no cross-cluster network path existing. It must be revisited (checking resolved cluster/trust-domain identity, not just suffix) the moment ADR-0001's multi-region rollout introduces any legitimate cross-cluster routing — see Forward Pointers. Until then, admission control and network policy remain in place as backstops rather than the primary line of defense.

Open Items

  • Device/tenant metadata for opted-in customer devices in QA. Does full device-record mirroring need to happen alongside telemetry mirroring, does QA need a lightweight shadow/stub device record, or does v1 scope down to BlueGuardian's own test-fleet devices only (already dual-provisioned) and defer real customer-device metadata sync to a follow-up? Not yet decided.

  • Exact retroactive-purge mechanism. What actually executes the QA-scoped purge on consent revocation, and how does it reuse or parallel whatever implements ADR-0005's four-state data lifecycle purge execution? Not yet designed.

  • Customer-facing opt-in surface. The actual settings UI/API for the customer-facing toggle (webApp, KMP shared model) is out of scope for this ADR and needs its own follow-up in the relevant frontend ADR track.

  • Bridge service ownership. Whether this lives in a new dedicated repo or is folded into an existing service's repo is not yet decided.

  • Topic naming for mirrored records. This decision reuses hono.telemetry.cerebral-stratum as the target topic name in QA (so mirrored and QA-native traffic are indistinguishable to consumers). Revisit if that intermixing turns out to cause confusion downstream — e.g. for anomaly-detection model training, which may want to exclude mirrored-from-prod records.

  • Kyverno vs. RHACS admission-control adoption. The regional-boundary enforcement's second layer depends on which (or both) is adopted; overlapping webhook design is already an open item at the platform level and this ADR inherits that dependency.

Forward Pointers

  • An implementation Epic will be filed in YouTrack (CSPROD project) once this ADR is accepted, covering: backend's device-attribute + opt-in API + Kafka-producer work, the new bridge service itself, and infra/gitops KafkaUser /topic wiring in both environments.

  • Regional-boundary enforcement must be upgraded, not just revisited, once/if ADR-0001's multi-region rollout is actually implemented. Specifically: the moment any legitimate cross-cluster network path exists (e.g. a Cloudflare tunnel for control-plane routing), the DNS-suffix check alone is no longer sufficient — it must be replaced or supplemented with a check on the resolved address's actual cluster/trust-domain identity. Flagged here specifically so it isn't silently missed at that point.

  • The retroactive-purge-on-revocation mechanism should be designed alongside, or reuse, whatever eventually implements ADR-0005's stated four-state purge lifecycle, rather than inventing a second purge path.

Last modified: 12 September 2026