Platform deployment architecture
The release distribution combines one shared platform definition with exactly one environment profile. Telemetry is an optional third layer; it is not a dependency of workflow execution.
flowchart LR
subgraph Host[Docker host]
subgraph Edge[edge network]
T[Traefik]
UI[Studio]
E[Engine replica or replicas]
K[Keycloak - development only]
end
subgraph Data[data network - internal]
DB[(PostgreSQL)]
KD[(Keycloak PostgreSQL - development only)]
end
subgraph Telemetry[telemetry network - internal and optional]
C[OpenTelemetry Collector]
P[Prometheus]
J[Jaeger]
L[Loki and Alloy]
G[Grafana]
end
end
B[Browser or API client] --> T
T --> UI
T --> E
T --> K
E <--> DB
K <--> KD
E -. bounded OTLP export .-> C
E -. named JSON log volume .-> L
C --> P
C --> J
P --> G
J --> G
L --> G
Layer ownership
Section titled “Layer ownership”| Layer | Owns | Must not own |
|---|---|---|
compose.yaml |
PostgreSQL, Engine, Studio, the Agent Worker, Docs, internal data and edge networks, persistent workflow/log volumes | Identity provider, TLS policy, telemetry backends, LLM credentials |
compose.dev.yaml |
Local HTTP routing, bundled Keycloak, safe evaluation credentials, automatic Agent Worker provisioning | Production identity or certificates |
compose.prod.yaml |
Exact application version, TLS/ACME routing, required external OIDC and CORS values, production restart/resources | Bundled production Keycloak or default secrets |
compose.telemetry.yaml |
Collector, metrics, traces, logs and Grafana | Database authority, engine readiness or workflow transactions |
The release bundle contains these files and every mounted configuration asset. It intentionally contains no Docker build context, so deployments pull the same immutable frontend and engine images everywhere.
Agent Worker and Insight placement
Section titled “Agent Worker and Insight placement”The Agent Worker is a platform service: it connects to the engine through
worker protocol v1, holds only the LLM endpoint/key and model allow-lists it
needs, and never exposes those to Studio or the engine. The Insight Engine is
an out-of-band analyzer inside the engine process (scheduled when
ABADA_INSIGHT_ENABLED=true) that reads durable fact windows from PostgreSQL
and optionally calls the configured LLM endpoint for proposal drafts. Both
keep model calls outside workflow transactions and both stay optional for the
certified core topology: a run never depends on a worker being present (it
pauses at agent work) and never depends on the analyzer.
Frontend startup contract
Section titled “Frontend startup contract”Studio does not embed installation-specific URLs during vite build.
Its container entrypoint validates API URL, OIDC URL, realm and client ID,
escapes them into /config.js, and starts Nginx only after validation passes.
The HTML loads /config.js before the application bundle, and Nginx marks that
file as non-cacheable. This lets one image serve multiple installations while
still failing fast on incomplete configuration.
Workflow and telemetry health
Section titled “Workflow and telemetry health”The engine Compose health check uses /api/actuator/health/readiness. Its
readiness group contains application and PostgreSQL state and excludes
telemetry. /api/actuator/health/telemetryExport reports whether export is
disabled or configured, but never probes or gates on a collector.
With telemetry disabled, the engine supplies a no-op Tracer to command
services and creates no span or metric exporter. With telemetry enabled, span
queues, batches, retry time and request timeouts are bounded. Export happens
outside workflow-state transactions; an outage may lose diagnostic signals
but cannot roll back a committed task or make readiness fail.
Persistent and replaceable state
Section titled “Persistent and replaceable state”postgres_datais authoritative workflow state and must be backed up.engine_logscarries structured rolling logs to Alloy when the overlay is present; logging continues without Alloy.letsencrypt_datapreserves production ACME state.- Prometheus, Loki and Grafana volumes are operational telemetry state. Their loss does not change workflow correctness.
- Engine and frontend containers are replaceable. Replica coordination and restart recovery come from PostgreSQL, not container-local memory.