Skip to content

Packaged API + UI Control Plane

You do not need to read this unless you are installing or customising the agent-bom Helm chart directly (instead of going through the reference installer in Deployment Overview).

agent-bom now ships a Helm-packaged control plane for teams that want the API and dashboard inside their own Kubernetes environment instead of running custom Deployment manifests by hand.

This is the right path when you want:

  • the API and UI in your own cluster
  • same-origin browser traffic through your own ingress
  • Postgres, ClickHouse, SSO, and secrets kept in your own environment
  • the scanner CronJob and optional runtime monitor packaged alongside the control plane
  • production operator defaults without pretending there is a managed vendor plane
  • a clean split between the API/runtime image and the standalone UI image

When you also need Terraform ownership for the AWS baseline outside the cluster, pair this chart with the Terraform AWS Baseline module. Terraform should own RDS, S3, IAM/IRSA, and Secrets Manager; Helm should own the in-cluster Deployments, CronJobs, and ExternalSecret objects.

The reverse path is now explicit too:

export AWS_REGION="<your-aws-region>"
agent-bom teardown \
  --cluster-name agent-bom-prod \
  --region "$AWS_REGION" \
  --namespace agent-bom \
  --release agent-bom \
  --dry-run

That helper tears down the chart first and the product-owned Terraform baseline second, while leaving platform-owned cluster infrastructure alone.

Chart removal now includes packaged Helm pre/post-delete hooks that clean up product-owned in-cluster leftovers such as generated ExternalSecret target secrets, CronJobs, Jobs, and PVCs before the Terraform baseline is destroyed.

What the chart deploys

When you set controlPlane.enabled=true, the Helm chart can package:

  • API Deployment + Service
  • UI Deployment + Service
  • same-origin Ingress that routes API paths to the API service and / to the UI
  • scanner CronJob
  • optional runtime monitor DaemonSet with a dedicated service account and no automounted service-account token by default

The image split is intentional:

  • agentbom/agent-bom runs the API, scanner jobs, gateway, proxy-related entrypoints, and other non-browser workloads
  • agentbom/agent-bom-ui runs the standalone browser UI that sits behind the same ingress or a separate UI service

Enterprise deployment topology

The canonical self-hosted topology and runtime/data-flow diagrams now live in Deployment Overview.

Use that page when you need to answer:

  • what runs where in customer-controlled infrastructure
  • which components are core vs optional per rollout
  • how scans, fleet, proxy, gateway, and exports flow back into the control plane

This page stays focused on the Helm-packaged control-plane shape itself:

  • what the chart deploys
  • how same-origin ingress is wired
  • what defaults are secure by design
  • how to install and operate the packaged API + UI control plane

Choose the deployment model before scaling

Model State and configuration Availability boundary
Single-instance VM or Kubernetes pilot One API instance; persistent local state; configured authentication. The shipped SQLite pilot is a demo profile with anonymous viewer access. Restart recovery depends on its volume; this is not an HA configuration.
On-premises Kubernetes Customer PostgreSQL, shared session/audit keys, ingress TLS, identity, and artifact storage. Use the BYO PostgreSQL secret layout and operator-owned values. API/UI replicas require node placement, database failover, and storage recovery to be exercised together.
Customer-cloud Kubernetes The same shared-state requirements; managed PostgreSQL, object storage, and cloud identity can supply those services. EKS examples include AWS-specific annotations and secret providers. A provider-compatible connection string is not evidence of tested provider failover or restoration.
Air-gapped deployment Apply offline images, packages, and vulnerability-data procedures to the selected single-instance or HA model. Disconnected operation does not supply database HA or shared storage; maintain an explicit data-refresh and upgrade process.

Render the selected operator values before installing:

helm template agent-bom deploy/helm/agent-bom \
  --namespace agent-bom --values ./my-prod-overrides.yaml > rendered.yaml

Review the rendered API environment, Secrets, migration Job, probes, ingress, volumes, and disruption budgets. Then exercise authentication, a scan, report download from another replica, and a controlled pod restart in the target cluster. Rendering validates configuration; it does not prove live failover. Use the shipped profile matrix and BYO PostgreSQL as configuration references. Preserve the separation between tenant-bound app, scoped maintenance, and migration/admin credentials.

Replica safety and disruption budgets

The chart derives AGENT_BOM_CONTROL_PLANE_REPLICAS from the largest configured API replica count: initial replicas, an enabled autoscaler's minimum/maximum, and its KEDA fallback. A deployment that starts at one but can grow to six therefore enables the application's shared-state safeguards from its first pod. This is a configuration safety signal, not a live pod-count metric. Set replicas and autoscaling through chart values; overriding this variable in controlPlane.api.env is rejected.

When pdb.enabled=true, the API/UI budget also accounts for autoscaling above one replica. A minimum of one replica can still block voluntary eviction of that sole healthy pod. Choose an HA floor of at least two and review placement across the failure domains available in the customer cluster. Disruption budgets constrain supported voluntary evictions; they do not prevent node failures or govern Deployment rolling updates. See Kubernetes disruptions.

PostgreSQL connection budget

Budget all pools in src/agent_bom/api/postgres_common.py, not only the main application pool. Each API process can open:

  • Application pool: max(1, AGENT_BOM_POSTGRES_POOL_MIN_SIZE, AGENT_BOM_POSTGRES_POOL_MAX_SIZE).
  • Idempotency fencing pool: max(1, min(4, AGENT_BOM_POSTGRES_POOL_MAX_SIZE)).
  • Maintenance pool: the same bound as the fencing pool, with separate credentials.

With the default minimum 5 and maximum 20, that is 28 direct connections per API process. Six replicas with one process each can use 168; the KEDA example's 12 replicas can use 336. Threads share a process's pools; extra API processes multiply this budget. Add gateway/other clients, migrations, backups, and operational reserve separately. Reserve capacity for rolling-update surge and terminating pods too; the steady-state scaler maximum is not a hard bound on all concurrently existing pods. See Kubernetes Deployment updates.

For a connection pooler, budget client slots and database connections separately and verify its pooling mode with tenant/RLS and transaction tests. Keep consistency-sensitive traffic on the writable PostgreSQL service endpoint. The customer database service owns primary election and failover; asynchronous standbys can lag and have different durability tradeoffs from synchronous replication. See PostgreSQL standby replication. Backup restoration and point-in-time recovery require a separate rehearsal.

All replicas serving report downloads need the same artifacts: configure AGENT_BOM_REPORT_S3_BUCKET, or a shared mounted directory with AGENT_BOM_REPORT_ARTIFACT_DIR and AGENT_BOM_REPORT_ARTIFACT_SHARED=1. A pod-local directory cannot supply cross-pod durability. Configure PVCs and storage classes for the customer's storage system; RAID belongs to that layer.

Gateway writers, partitions, and shards

The gateway remains one PVC-backed writer. Its delivery claims and local audit state are process-owned; the chart rejects multiple replicas and gateway autoscaling. Shared delivery leases, stale-worker fencing, idempotent retry, and audit-order recovery must be implemented and tested before that guard can be lifted.

PostgreSQL audit-table partitioning divides tables inside one database; it does not spread writes across independent database primaries. Automatic database sharding is not implemented. Measure concurrent tenants, ingest, graph growth, and recovery on a correctly sized primary first. If that becomes the measured limit, tenant placement across database clusters needs explicit migration, schema-upgrade, and cross-tenant query contracts.

Optional ClickHouse analytics

Keep transactional control-plane state in PostgreSQL for shared-state API deployments, or SQLite for a single instance. The existing ClickHouse backend stores analytics such as vulnerability trends, runtime events, and posture history. Self-hosted ClickHouse can provide that destination; it does not replace the PostgreSQL jobs, authorization, and graph stores.

With the control-plane database already configured, inject AGENT_BOM_CLICKHOUSE_USER and AGENT_BOM_CLICKHOUSE_ACCESS_TOKEN through the deployment's secret mechanism, then start the API with its ClickHouse HTTP endpoint:

agent-bom serve --analytics-backend clickhouse \
  --clickhouse-url https://clickhouse.example.com:8443 --no-ui

In Helm, set AGENT_BOM_ANALYTICS_BACKEND=clickhouse and AGENT_BOM_CLICKHOUSE_URL in controlPlane.api.env, and inject credentials from the existing Secret. After recording scan analytics, query the tenant's history with agent-bom report analytics trends --tenant your-tenant using the same ClickHouse connection settings. Check the analytics tables and API logs before enabling higher ingest volume.

Buffered analytics uses a bounded queue and can drop evidence when that queue overflows; it is not the authoritative audit ledger. ClickHouse replication and shards require separate database configuration and qualification. A provider's managed PostgreSQL service is a separate PostgreSQL endpoint and must pass the same migrations, permissions, and tenant-isolation checks as other PostgreSQL deployments.

Before the first production scan

Refresh vulnerability intelligence inside the customer-controlled environment before running the first scheduled or offline scan:

agent-bom db update
agent-bom db freshness --format json

The freshness result is evidence for the scan run, not a guarantee that every upstream advisory is available. Keep the refresh job inside the same network and change-control boundary as the control plane; for air-gapped clusters, import a reviewed database artifact and record its source timestamp. The already-shipped scripts/pilot-verify.sh then provides the non-destructive health, auth, scan, fleet, and signed-evidence smoke check for a deployed instance.

For the runtime operator guides behind this flow, see:

Same-origin default

The UI runtime contract from #1452 is what makes this honest.

By default the chart leaves NEXT_PUBLIC_API_URL blank in the UI pod, so the browser uses relative paths:

  • /v1/*
  • /healthz (and /health — same JSON liveness probe)
  • /docs
  • /redoc
  • /openapi.json
  • /ws/*

The packaged ingress routes those paths to the API service and everything else to the UI service. That means:

  • one hostname
  • no CORS setup for the default path
  • no UI image rebuild per environment

If you want cross-origin instead, set controlPlane.ui.env so NEXT_PUBLIC_API_URL points at the API host you own.

Secure-by-default boundaries

The chart packages the control plane, but it does not quietly weaken the runtime model.

  • API and UI pods run with automountServiceAccountToken: false
  • the optional monitor DaemonSet now also runs with automountServiceAccountToken: false and its own service-account path instead of inheriting the shared chart identity
  • the discovery service account and IRSA path stay attached to the scanner
  • the API still refuses non-loopback startup without AGENT_BOM_API_KEY, OIDC, SAML-issued session keys, or an explicit insecure override
  • multi-replica API deployments now require a PostgreSQL-backed shared rate-limit store; the chart exports the replica floor and the API fails closed instead of silently falling back to process-local limits
  • same-origin ingress avoids default CORS sprawl
  • network policy stays enabled, with configurable ingress restrictions

Minimal values example

Create distinct database and auth Secrets (the shipped example contains all four object shapes and placeholder keys):

cp deploy/helm/agent-bom/examples/postgres-secret.example.yaml /tmp/agent-bom-postgres-secrets.yaml
# Replace every placeholder from the platform secret manager, then:
kubectl apply -f /tmp/agent-bom-postgres-secrets.yaml
kubectl create secret generic agent-bom-control-plane-auth -n agent-bom \
  --from-literal=AGENT_BOM_API_KEY='replace-me'

Then install with a values file like:

controlPlane:
  enabled: true
  postgresSecrets:
    enabled: true
    appSecretRef:
      name: agent-bom-control-plane-db
    maintenanceSecretRef:
      name: agent-bom-control-plane-maintenance
    adminSecretRef:
      name: agent-bom-control-plane-admin
  api:
    envFrom:
      - secretRef:
          name: agent-bom-control-plane-auth
  ingress:
    enabled: true
    className: nginx
    hosts:
      - host: agent-bom.internal.example.com

scanner:
  enabled: true
  allNamespaces: true
  serviceAccount:
    annotations:
      eks.amazonaws.com/role-arn: arn:aws:iam::REPLACE_ME_ACCOUNT_ID:role/REPLACE_ME_AGENT_BOM_DISCOVERY_ROLE

serviceAccount:
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::REPLACE_ME_ACCOUNT_ID:role/REPLACE_ME_AGENT_BOM_DISCOVERY_ROLE

The chart supports component-specific service-account overrides for scanner, gateway, and backup jobs. If you omit the component-specific annotations, they inherit the shared serviceAccount.annotations block.

Install:

helm install agent-bom oci://ghcr.io/msaad00/charts/agent-bom \
  --version 0.108.4 \
  -n agent-bom --create-namespace \
  -f values.agent-bom.yaml

Verify the rollout before exposing it

Run these checks from an operator workstation after the Helm install. They exercise the same liveness, readiness, and authenticated API paths used by the Ingress and Kubernetes probes; a green Helm command alone does not prove that the control plane is ready to receive evidence.

kubectl -n agent-bom rollout status deployment/agent-bom-api --timeout=5m
kubectl -n agent-bom rollout status deployment/agent-bom-ui --timeout=5m

# Keep the API private while validating it.
kubectl -n agent-bom port-forward svc/agent-bom-api 18080:8000 >/tmp/agent-bom-port-forward.log 2>&1 &
PF_PID=$!
trap 'kill "$PF_PID" 2>/dev/null || true' EXIT
curl --fail --silent http://127.0.0.1:18080/healthz
curl --fail --silent http://127.0.0.1:18080/readyz
curl --fail --silent -H "Authorization: Bearer $AGENT_BOM_API_KEY" \
  http://127.0.0.1:18080/v1/auth/me

/healthz is liveness only; /readyz must be healthy before routing traffic, and /v1/status confirms that the credential and tenant/auth configuration are usable. If readiness fails, keep the Ingress private, inspect kubectl -n agent-bom logs deploy/agent-bom-api, and fix the backing database or secret reference before retrying. Do not set an insecure-auth override just to make a probe pass.

For a reversible release, record the current revision and use Helm rollback if the new deployment fails its rollout or readiness checks:

helm -n agent-bom history agent-bom
helm -n agent-bom rollback agent-bom <known-good-revision> --wait --timeout 5m

Single-node SQLite pilot preset

For a fast in-cluster pilot without external Postgres, use the shipped single-node preset:

Create the auth secret and a small PVC first:

kubectl create secret generic agent-bom-control-plane \
  -n agent-bom \
  --from-literal=AGENT_BOM_API_KEY='replace-me'

kubectl apply -n agent-bom -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: agent-bom-sqlite-pilot
spec:
  accessModes: ["ReadWriteOnce"]
  resources:
    requests:
      storage: 10Gi
EOF

Then install:

helm install agent-bom oci://ghcr.io/msaad00/charts/agent-bom \
  --version 0.108.4 \
  -n agent-bom --create-namespace \
  -f deploy/helm/agent-bom/examples/eks-control-plane-sqlite-pilot-values.yaml

This preset is intentionally single-node only: one API pod, one UI pod, PVC-backed SQLite state, no HPA, no PDB, and no multi-replica availability claims.

Production defaults example

For the stronger self-hosted operator path, start from:

That example adds:

  • HPA for API and UI
  • ServiceMonitor enabled in the production preset so /metrics is scraped when Prometheus Operator is present
  • HPA scale-down stabilization
  • topology spread across zones and nodes
  • preferred pod anti-affinity for API and UI replicas
  • optional control-plane PriorityClass
  • fail-closed shared rate limiting when the Postgres-backed limiter is unavailable
  • cert-manager ingress annotations and TLS wiring
  • external-secrets integration for the control-plane secrets, including split refresh cadence for DB vs auth/HMAC material
  • packaged PrometheusRule alerts for API error rate, scanner failures, OIDC decode failures, and proxy audit backlog
  • packaged Grafana dashboard ConfigMap for clusters that already watch dashboard config
  • packaged Postgres backup CronJob that runs pg_dump and uploads to S3 through IRSA with SSE or KMS
  • dedicated service-account hooks for gateway and backup jobs, inheriting the scanner IRSA annotations unless you override them
  • restricted ingress defaults for the chart network policy
  • optional cert-manager-backed sidecar auto-injection webhook for HTTP/SSE MCP workloads

For clusters that already standardize on a service mesh and policy controller, start from:

That example adds:

  • packaged Istio PeerAuthentication for strict mTLS on agent-bom pods
  • packaged Istio AuthorizationPolicy that keeps same-namespace traffic and explicitly whitelisted ingress namespaces
  • packaged namespaced Kyverno Policy that enforces the same restricted pod contract already used by the chart

This is intentionally an opt-in hardening layer. It composes with the chart's existing NetworkPolicy, PSS-restricted pod settings, anti-affinity, and HPA defaults instead of replacing them.

What you still own

This is a real packaged control plane, but not a magic managed service.

You still own:

  • Postgres and optional ClickHouse
  • ingress controller and TLS
  • OIDC or SAML IdP configuration, or API key secret management
  • cluster-specific autoscaling thresholds and failure-domain policy
  • operator runbooks and load testing

Production guidance

  • keep controlPlane.api.replicas and controlPlane.ui.replicas at 2+
  • use Postgres, not SQLite
  • database migrations run automatically: the chart ships a pre-install,pre-upgrade Helm hook (controlPlane.migrations.enabled, on by default) that runs alembic -c deploy/supabase/postgres/alembic.ini upgrade head before the new API pods roll. No manual Alembic step is required for Helm upgrades.
  • databases first bootstrapped from init.sql should be stamped once with the baseline so the auto-hook has a revision to upgrade from: alembic -c deploy/supabase/postgres/alembic.ini stamp 20260416_01
  • to manage migrations yourself (non-Helm or external DB tooling), set controlPlane.migrations.enabled=false and run upgrade head in your pipeline
  • enable the control-plane HPAs before higher-volume rollout
  • use anti-affinity and a control-plane PriorityClass when you expect node pressure
  • enable topology spread when you run multi-AZ EKS
  • keep same-origin ingress unless you have a strong reason not to
  • use distinct envFrom / Secrets for app, maintenance, and migration/admin Postgres credentials; use the auth Secret for API keys, OIDC issuer, audience, optional required nonce, SAML IdP/SP metadata values, and audit HMAC settings
  • configure the three distinct controlPlane.postgresSecrets references. The chart projects app + maintenance to the API, admin + app + maintenance to the idempotent migration/reconcile Job, and maintenance only to backups
  • enforce API key lifetime policy with AGENT_BOM_API_KEY_DEFAULT_TTL_SECONDS and AGENT_BOM_API_KEY_MAX_TTL_SECONDS; admin key replacement uses POST /v1/auth/keys/{key_id}/rotate so rotation stays explicit and audited
  • split fast-rotating auth secrets from slower DB config with controlPlane.externalSecrets.secrets[] so AGENT_BOM_OIDC_*, AGENT_BOM_SAML_*, and AGENT_BOM_AUDIT_HMAC_KEY can refresh at 5m while the app and maintenance Postgres credentials stay at 1h
  • set AGENT_BOM_REQUIRE_SHARED_RATE_LIMIT=1 for multi-replica production control planes so the API refuses to start if the shared limiter backend is unavailable
  • tune Postgres-backed control planes explicitly with AGENT_BOM_POSTGRES_POOL_MIN_SIZE, AGENT_BOM_POSTGRES_POOL_MAX_SIZE, AGENT_BOM_POSTGRES_CONNECT_TIMEOUT_SECONDS, and AGENT_BOM_POSTGRES_STATEMENT_TIMEOUT_MS
  • enable controlPlane.observability.prometheusRule.enabled=true when the cluster already runs Prometheus Operator
  • keep monitor.enabled=true and monitor.serviceMonitor.enabled=true in the production preset unless your platform team has a different scrape contract
  • enable controlPlane.observability.grafanaDashboard.enabled=true when Grafana watches dashboard ConfigMaps
  • enable controlPlane.backup.enabled=true only after setting a real S3 bucket, prefix, IRSA-backed upload permissions, and a maintenance-only canonical maintenance Secret containing AGENT_BOM_POSTGRES_MAINTENANCE_URL
  • enable controlPlane.serviceMesh.enabled=true only when the control-plane namespace is already part of your Istio data plane
  • keep controlPlane.serviceMesh.istio.authorizationPolicy.allowedNamespaces explicit; the packaged example allows ingress-nginx and istio-system, but production should match your real ingress path
  • enable controlPlane.policyController.enabled=true only when Kyverno is already installed cluster-wide; the chart packages the namespaced policy, not the controller itself
  • set controlPlane.backup.destination.bucketRegion to the actual region of your backup bucket; the production example intentionally uses REPLACE_ME_BUCKET_REGION
  • controlPlane.backup.destination.region remains as a backward-compatible fallback for older values files
  • keep controlPlane.backup.destination.encryption.enabled=true; the default is AES256, and production should set mode=aws:kms with a dedicated kmsKeyId
  • restore drills should use deploy/ops/restore-postgres-backup.sh: ./deploy/ops/restore-postgres-backup.sh s3://bucket/key.dump "$RESTORE_POSTGRES_URL" REPLACE_ME_BUCKET_REGION
  • RESTORE_POSTGRES_URL must use the migration/admin principal because the restore path performs destructive DDL; never expose it to API or backup pods
  • use Backup and Restore Runbook for the full operator checklist
  • expose GET /v1/auth/saml/metadata to your IdP admins and keep POST /v1/auth/saml/login behind the same ingress hostname as the API
  • enable PDBs when you are running multi-replica workloads

Current boundary

The chart now packages the control plane honestly, but it still does not claim:

  • a bundled Postgres subchart
  • benchmarked throughput guarantees
  • completed auth hardening beyond the currently shipped server contract

Those are the next operator-hardening layers, not hidden assumptions.