CEREBRAL STRATUM Help

ADR-0006: Keycloak Deployment Topology — RHBK Operator, Per-Cluster Instances Federated with Production

Field

Value

Status

Proposed

Date

2026-09-07

Author

Alex Henshaw

Supersedes

ADR-0003 (Identity Infrastructure Placement — Keycloak on ECS with RDS)

Relates to

cerebralstratum ADR-0001 (Network Topology), ADR-0004 (Admin-Plane Access Control), ADR-0005 (Data Classification); cerebralstratum-backend ADR-0001 (IAM & Device Registration), ADR-0005 (UMA 2.0 Device Resources), ADR-0007 (Emergency SOS), ADR-0009 (Law Enforcement Device Share)

Tracked in

CSPROD-253, CSPROD-269

See also

YouTrack CSPROD-A-38 ("ADR-0008: Secrets Management" — drafted, not yet migrated into this repo)

Context

ADR-0003 proposed migrating Keycloak to ECS Fargate + RDS as a bootstrap-budget interim step, justified by ROSA HCP being cost-decommissioned at the time. That migration never happened. Actual deployment took the path backend ADR-0001 always described — the RHBK Operator on OpenShift — but arrived at that point without a governing platform-wide ADR, and the current state has real gaps:

  • Only one Keycloak instance exists in git (infra/3-operators/rhsso), with no database wired to it at all (Keycloak CR has no spec.db), no realm codified anywhere, and its overlay structure fans out to spoke clusters via ACM placement rather than the hub — none of which matches ADR-0001's hub-and-spoke model (identity is a hub concern).

  • Per CSPROD-253/CSPROD-269, QA currently has no Keycloak instance or realm at all — cerebral-stratum-backend cannot reach Running in any environment until this is resolved.

  • ACM was temporarily removed from the estate (CSPROD-259/261), so the spoke fan-out isn't currently live, but the underlying manifests would resume that behaviour once ACM/spoke scaling returns, unless corrected now.

This ADR records the actual target architecture and closes the gap left by ADR-0003's unimplemented ECS path.

Decision

RHBK Operator, deployed per environment on the core hub cluster, one Keycloak instance per environment, federated with production for admin access only.

Topology

  • One RHBK-operator-managed Keycloak CR per environment (QA, prod), each in its own namespace (sso, sso-qa) on the core hub cluster — consistent with ADR-0001's hub-and-spoke model (identity is a hub concern; spokes host tenant device-telemetry only) and with how core already hosts both prod-hub and QA-app-workload namespaces during the current single-cluster bootstrap phase.

  • Each instance backed by its own Crunchy PostgresCluster, following the same operator-generated-credentials pattern already used for the app's own database (gitops/.../supporting-infra/database) — no spec.db-less instances going forward.

  • QA is built first as the testbed; production Keycloak (identity.blueguardian.co) already exists and remains the federation anchor. (Renamed from sso.blueguardian.co — that hostname is cluster/infra's own SSO, realms/internal, not this workload-facing realms/external instance; see CSPROD-285.)

  • Reconciling "hub concern" with future per-spoke instances: ADR-0001's "identity is a hub concern" statement describes the current topology, where the hub is a single administrative/control point and spokes hold no regionally-bound customer data of their own. It is not a permanent architectural ban on any Keycloak instance ever running on a spoke. When spoke-per-region scaling resumes, each spoke becomes the regional home for that region's own customer data (per cerebralstratum ADR-0005's data-sovereignty boundary — regionally-bound data, including session state and token contents, must not cross region boundaries even transiently). At that point, a region's own users authenticating against that region's own Keycloak instance is the sovereignty-correct design, not a violation of ADR-0001 — the spoke is acting as the regional identity authority for its own tenants, the same role the current single hub plays for the whole (currently single-region) platform. What ADR-0001 rules out is a spoke depending on a different region's Keycloak for its own users' auth, or platform-admin/control-plane identity living anywhere but the hub — neither of which this ADR proposes. This is why the topology in this ADR is deliberately framed as "per-cluster" rather than "per-environment": it already anticipates the scaled-out shape, even though only core is active today.

Repository split

  • RHBK operator subscription/install → infra repo (3-operators/rhsso), matching the existing convention for crunchydata, amq-streams, etc.

  • Keycloak CR instances, per environment → gitops repo, as a new supporting-infra/keycloak/ subsystem sibling to database/message-bus/misc under each environment's cerebral-stratum workload tree.

Realm as code

cerebralstratum-backend's backend/src/main/resources/devservices/realm.json is left unchanged — it remains exactly what it is today, the Quarkus Dev Services realm for local development. It is the structural starting point for QA/prod realm definitions (roles, device-fleet's UMA authorizationSettings, cerebral-stratum-frontend, platform-admins), not their source of truth going forward — the two are deliberately allowed to diverge from this point on.

The actual source of truth is a cleaned, per-environment realm config (credentials and dev-only test fixtures stripped/replaced) committed to the gitops repo under the new supporting-infra/keycloak/ subsystem, rendered as a ConfigMap via Kustomize alongside each environment's Keycloak CR.

Sync mechanism: keycloak-config-cli as a sidecar in the Keycloak pod, injected via the RHBK Operator's spec.unsupported.podTemplate (a real, documented mechanism for adding extra containers to the operator-managed pod — Red Hat's own designation as "unsupported" means best-effort/no compatibility guarantee across operator versions, not that it doesn't work). Since keycloak-config-cli itself is a one-shot idempotent importer with no built-in watch/interval mode, the sidecar wraps it in a simple loop (run → sleep → run again) rather than relying on a feature the tool doesn't have. The mounted ConfigMap picks up git changes through the normal ArgoCD sync → kubelet volume-refresh path (~1 minute typical propagation); the sidecar's next loop iteration reconciles the live realm against whatever is currently mounted.

Security note carried from the podTemplate mechanism itself: the sidecar has access to the same namespace Secrets as the main Keycloak container, including its own admin API credentials. Those credentials should be operator/automation-generated per the secrets direction described in YouTrack CSPROD-A-38 ("ADR-0008: Secrets Management" — drafted but not yet migrated into this repo's Writerside/topics/ADRs/; referenced here as the current decision record, not as an in-repo ADR) — the same open problem as backend-oidc (CSPROD-253), not a new one.

Backend ADR-0007's ephemeral emergency-contact identities and the dev realm's IT-test-fixture users are both realistic shapes worth adapting into QA's realm (not the literal dev credentials) so QA exercises the same identity-growth pattern production will see.

Federation — admin convenience only, explicitly bounded away from device data

Each cluster's Keycloak federates with production Keycloak (identity.blueguardian.co) as an identity broker, so platform employees administer every instance from one identity instead of separate per-cluster admin accounts. This federation is scoped strictly to admin-plane reachability and FGAP V2-scoped permissions (ADR-0004) — logging into a given cluster's admin console as a federated prod identity.

Explicit boundary, required by backend ADR-0009: federation must never become a path by which a federated admin identity can originate or hold a UMA 2.0 device-share grant (device:share) on that cluster's realm. ADR-0009 is unconditional that only the device owner's own authenticated subject may create an LE-share permission, and backend ADR-0005 already prohibits any standing platform-admin policy on device:read/device:modify. A federated admin identity is exactly the kind of identity ADR-0009 was written to exclude from having any device-data path — federation config (the IdP broker's default roles/scope mappings) must not grant one implicitly. This needs explicit verification once federation is configured, not an assumption that FGAP V2 scoping alone handles it.

Secrets

Following the direction already established for Kafka/Postgres credentials (CSPROD-244) and described in YouTrack CSPROD-A-38 ("ADR-0008: Secrets Management" — not yet migrated into this repo): backend-oidc moves away from 1Password-sourced secrets toward operator/automation-generated secrets. Exact mechanism (RHBK-native client secret generation vs. a bootstrap Job) is open scope, tracked in CSPROD-253.

Alternatives Considered

  • Continue ADR-0003's ECS Fargate + RDS path. Rejected — never implemented, and contradicts backend ADR-0001's OpenShift-native description; would require migrating already-working RHBK-on-OpenShift infrastructure to AWS for no operational gain.

  • Single shared Keycloak instance/realm across QA and prod. Rejected — no environment isolation for identity while every other piece of the stack (Postgres, Kafka, Redis) is already properly separated per environment; a QA realm change or load test could affect production auth.

  • Fully isolated per-environment Keycloak, no federation. Considered as the simpler, stricter-isolation option. Rejected in favour of federation for admin convenience — with more than one cluster, unfederated per-cluster admin accounts create real operational friction (separate credentials, separate MFA enrolment, no single audit trail for admin actions) for a cost (federation complexity) that's manageable given the explicit boundary against device-data access above.

  • RHBK KeycloakRealmImport CR for realm sync. Rejected in favour of keycloak-config-cli — the operator's realm import CR is closer to an apply-once import than a continuously reconciled sync, and doesn't fit "the realm definition lives in git and is kept in sync" as an ongoing property.

  • kcadm.sh (Keycloak's bundled admin CLI) for realm sync. Considered for consistency with the rest of the Red-Hat-aligned stack. Rejected in favour of keycloak-config-cli — kcadm.sh is imperative (individual admin operations), not a declarative/idempotent config-to-live-state reconciler; keycloak-config-cli is purpose-built for exactly this git-realm-as-source-of-truth pattern.

Consequences

Positive

  • Closes the "QA has no Keycloak at all" gap blocking cerebral-stratum-backend from running in QA.

  • Identity infrastructure finally matches the isolation level already present elsewhere in the stack.

  • Realm content becomes reviewable, reproducible, and disaster-recoverable instead of living only in the admin console.

  • One admin identity across clusters reduces credential sprawl for platform employees.

Negative / Trade-offs

  • Federation is a new mechanism with a real, non-obvious failure mode (admin-plane access bleeding into device-data access) that must be explicitly verified, not assumed safe by construction.

  • Per-cluster Postgres instances for Keycloak add operational surface (backups, upgrades) multiplied by cluster count, though this mirrors an already-accepted pattern for the app DB.

  • The keycloak-config-cli sidecar runs under RHBK's "unsupported" podTemplate API surface — no compatibility guarantee across operator upgrades, needs re-verification whenever the RHBK operator version changes.

  • The sidecar's polling-loop reconciliation (not event-driven) means a small window between a git change landing and the live realm reflecting it, bounded by kubelet's ConfigMap refresh interval plus the sidecar's own sleep interval.

Open Items

  • Exact backend-oidc secret automation mechanism (CSPROD-253).

  • Verification procedure for the federation admin-plane/device-data boundary before QA federation goes live.

  • Whether rhsso's existing spoke-overlay wiring is corrected now or only when ACM/spoke scaling actually resumes.

  • LDAP/IdM (FreeIPA, ipa-primary) federation — still active per CSPROD-173, not yet reconciled with this per-cluster model; needs an explicit statement of whether each cluster's Keycloak federates to IdM independently or only prod does.

  • Concrete sleep interval for the keycloak-config-cli sidecar loop — not yet chosen.

  • keycloak-config-cli sidecar's own admin credential provisioning mechanism — tied to the same open backend-oidc-style secret automation question above.

Forward Pointers

  • Implementation Epic in YouTrack: CSPROD-253 (per-cluster RHBK deployment), with CSPROD-269 (missing backend-oidc secret) as a direct downstream dependent.

  • Secret automation design once YouTrack CSPROD-A-38 ("ADR-0008: Secrets Management") is migrated into this repo — backend-oidc and the keycloak-config-cli sidecar's own admin credentials should both be specified against that ADR's mechanism, not designed ad hoc here.

  • Sidecar sync verification and testing on the QA cluster deployment, including the federation admin-plane/device-data boundary check called out under Federation above — both should happen before QA federation is considered production-ready, and before this pattern is replicated to prod or future spoke clusters.

  • When ACM/spoke scaling resumes: revisit rhsso's existing spoke-overlay wiring (Open Items) and confirm the regional-Keycloak-per-spoke shape described under Topology against whatever the actual first spoke reprovisioning looks like.

Last modified: 16 September 2026