Architecture/Provisioning pipeline & the nexus fleet
Architecture

Provisioning pipeline & the nexus fleet

How the console's fleet (platforms, clusters, workspaces, tenants) is projected into nexus over EventBridge + SQS, how the snapshot backfill works, and the read-only fleet API nexus exposes.

This is an internal implementation reference for the console → nexus integration, not a public product guide. It corresponds to the outbox-events, provisioning-pipeline, cplane-deploys, and fleet-console repo skills.

Overview#

The console owns the fleet (platforms, clusters, workspaces, tenants). Nexus keeps a read-only projection of it for routing and billing. The console never calls nexus; it publishes domain events, and nexus consumes them.

text
console (Postgres tx + outbox)
  └─ EventBridge bus "console-prod"
       ├─ route_platform / route_global  → core regional buses   (actions: created|updated|deleted)
       └─ nexus-<domain> rules            → SQS "nexus-<domain>"  (all actions, incl. snapshot)
                                              └─ nexus inbox → Oban → Fleet/Workspaces/Tenants

Producing events (console)#

Writes are paired with a row in the transactional outbox (modules/outbox). A poller/ signal drains it, and processOutbox maps each row to an EventBridge detail-type:

Module
detail-type
platforms
platform.v1.<action>
clusters
platform.<public_name>.cluster.v1.<action>
workspaces
platform.<name>.cluster.<name>.workspace.v1.<action>
tenants
tenant.v1.<action>
auth_domains
auth_domain.v1.<action>

<action> is created, updated, deleted, or `snapshot` (the backfill — see below).

Routing (EventBridge)#

Two independent sets of rules sit on console-prod:

Core-bound (route_platform, route_global, owned by terraform/console) forward to

core's regional buses and are action-whitelisted to {created, updated, deleted}. snapshot is deliberately excluded, so a fleet backfill never reaches core.

Nexus-bound (nexus-<domain>, owned by terraform/cplane/modules/nexus) match by

detail-type only and forward to per-domain SQS queues. Nexus therefore receives snapshot.

Consuming events (nexus)#

text
SQS "nexus-<domain>" → Inbox.Consumer → Inbox.Store.ingest (persist + Oban)
  → ProvisioningWorker → Handlers.<Domain> → Fleet/Workspaces/Tenants upsert

Handlers upsert (idempotent, safe under SQS at-least-once and re-seeding) and whitelist the acting: created|updated|snapshot upsert, deleted is special, anything else is logged + skipped. Parents are resolved by `public_name` carried in the event payload (data.platform.public_name, data.cluster.public_name); platforms/clusters key on public_name, workspaces/tenants on id. A workspace also carries tenant_id (data.tenant.id) for tenant filtering.

Snapshot backfill (Seed)#

Because nexus is greenfield, POST /v1/<resource>/seed on the console re-emits every existing row as a snapshot event. Seeding order doesn't matter (nexus auto-creates missing parents), soft-deleted rows are correctly skipped, and it's safe to re-run.

Nexus fleet API (read-only)#

Synced from the console; _limit caps at 100 and every list has a matching /count for accurate totals.

Endpoint
Notes
GET /v1/platforms, /platforms/count, /platforms/:id
GET /v1/clusters, /clusters/count, /clusters/:id
filter platform_id
GET /v1/workspaces, /workspaces/count
paginated (_limit/_skip), _search, filters tenant_id/platform_id/cluster_id/status
GET /v1/tenants, /tenants/count, /tenants/:id

Failure modes worth knowing#

A new nexus domain must be added to Inbox.Message's domain allow-list, or messages are

delivered by EventBridge but dropped by the consumer — looks like a routing failure but isn't. Confirm delivery with CloudWatch AWS/Events Invocations/FailedInvocations.

Console outbox rows must be read with data::text (jsonb aliasing bug) or they intermittently

fail to publish and stall.