Phases 1–7 in review · not yet live-verified

One-box traffic shifting — a weighted router in front of production

2026-08-21 · ForgeGraph @ main · builds on deployment_targets · execution_policies (canaryWindowSeconds / autoPromoteAfterMs) · app_resources · @xyflow/react (topology page)

A green build is necessary but not sufficient. ForgeGraph's release path is Git repository → Build + unit-test evidence → Beta deployment + target verification → One-box live-traffic bake → Production. This plan adds an optional, per-app one-box: a second deployment target on the production stage that shares production's databases and secrets, sits behind a ForgeGraph-managed router, and receives a configurable slice of live traffic (default 10%) for a configurable bake window before the fleet is promoted. Steady state is 0% one-box / 100% production. A first-class application pipeline owns the ordered stages, gates, artifacts, verification suites, database boundaries and live traffic state.

10%default one-box slice during a deploy
0/100steady-state weights
2new tables: traffic_splits · traffic_events
0new stages — one-box is a deployment_target

Decision: how traffic is split

OptionSeparate Workers?Weight change w/o redeploy?Cost / fitVerdict
Router Worker (ours) — public hostname bound to <slug>-router, which service-binds <slug> and <slug>-onebox and picks by weightyesyes (KV / Worker settings)free tier; one extra hop (~1ms, same colo via service binding); works for node-platform apps too via fetch(origin)chosen
Workers Gradual Deployments (wrangler versions deploy --percentage)no — two versions of one Workeryesnative, freerejected: violates "separate Workers" requirement; can't mix platforms; no bake state visible to us
Cloudflare Load Balancer (pools/origins, weighted steering)yesyespaid add-on per zone; origin-oriented, Workers aren't first-class origins; LB session affinity is cookie-based anywayrejected: cost + poor Worker fit; revisit only for multi-node origin pools
The router is the load balancer feature. It's opt-in per app (traffic.enabled), and when disabled the public hostname binds straight to the production Worker exactly as today — nothing changes for apps that don't turn it on.

Goals & non-goals

One-box traffic shifting — weighted router in front of production

90% (idle: 100%)

10% (idle: 0%)

binding

binding

binding

binding

app.example.com
Workers custom domain

<slug>-router
weight from KV

<slug>
primary Worker

<slug>-onebox
one-box Worker

Postgres via Hyperdrive
fg-slug-production

D1 / R2 / KV
app_resources @ production

CONFIG KV
split:appId

UI slider / fg traffic set

Traffic topology with one-box enabled. Both production Workers bind the same stage-scoped DB and resources.
mermaid source
flowchart LR
  D[app.example.com<br/>Workers custom domain] --> R{{"&lt;slug&gt;-router<br/>weight from KV"}}
  R -- "90% (idle: 100%)" --> P["&lt;slug&gt;<br/>primary Worker"]
  R -- "10% (idle: 0%)" --> O["&lt;slug&gt;-onebox<br/>one-box Worker"]
  P -.binding.-> DB[(Postgres via Hyperdrive<br/>fg-slug-production)]
  O -.binding.-> DB
  P -.binding.-> RES[(D1 / R2 / KV<br/>app_resources @ production)]
  O -.binding.-> RES
  KV[(CONFIG KV<br/>split:appId)] --> R
  UI[UI slider / fg traffic set] --> KV

What exists today (and what this replaces)

ThingStateIn this plan
packages/api/src/routers/blue-green.tshard-disabled stub — blueGreenRuntimeSlotsAvailable() returns false, every mutation throws NOT_IMPLEMENTEDsuperseded; delete once traffic router lands
fg deploy promote → api/fg/deploy/[id]/promoteno-op (writes active back onto an active row)becomes "promote the one-box build to the primary target" (Phase 2)
agent/internal/deploy/canary.go MonitorCanary; PendingDeployment.CanaryWindowSeconds/RollbackOnFailuredead code — server never populates the fieldsbake moves hub-side (survives agent restarts, one place for the verdict); agent fields stay unused
execution_policies.canaryWindowSeconds / rollbackOnFailure / autoPromoteAfterMs, prod_canary lane, canary_paused intervention, lane-enforcement.tsschema + gating exist; nothing deploys to a canaryreused as-is; one-box is what prod_canary means
Workers custom domains (ensureWorkersDomain, cloudflare-client.ts:359)attached out-of-band via /api/fg/dns/workers-domains, never by deploy; errors if the hostname is bound to another scriptenable/disable = deleteWorkersDomain then ensureWorkersDomain for the router (one short window; done once, not per deploy)
Default worker name (cloudflare_worker.go:526)targetConfig.workerName → <slug>-<service> → <slug>; no stage componentone-box target always sets workerName explicitly
deployments partial unique indexes (deployment.ts:99-119)one in-flight + one active per (app, stage, target, node, service)one-box and primary are separate targets, so both can be active concurrently — no index change
Pipeline tabhand-rolled CSS-grid matrix over pipeline.current; React Flow lives in packages/ui/src/fleet-topology.tsxnew graph query + a React Flow pipeline component in packages/ui, matrix kept as the per-changeset history below it

Model

One-box is a second target on the production stage, not a fifth stage

Secrets, DB bindings and app_resources are scoped per (app, stage). Putting one-box on the production stage as a deployment_targets row (name: "one-box", platform: "cloudflare-workers", config.workerName: "<slug>-onebox") gives "shared databases, separate Workers" for free: the agent's existing wrangler patching (patchWranglerVars / planned patchWranglerBindings) emits the same Hyperdrive/D1/R2/KV bindings into both Workers because they resolve from the same stage. It also keeps STAGES.md's four canonical names intact and finally gives the legacy prod_canary lane (execution_requests.environment, lane-enforcement.ts, readiness.ts) a concrete thing it refers to: prod_canary ≡ production / one-box target, prod_main ≡ production / primary target.

New table traffic_splits (one row per app + stage, effectively production)

id, app_id, stage_id                     -- unique (app_id, stage_id)
enabled            boolean  default false
router_worker_name text                    -- "<slug>-router"
primary_target_id  text → deployment_targets
canary_target_id   text → deployment_targets   -- the one-box
canary_weight_bps  integer  default 0      -- live weight, basis points (1000 = 10%)
default_canary_bps integer  default 1000   -- what a deploy shifts to
bake_seconds       integer  default 900    -- seeded from execution_policies.canaryWindowSeconds
stickiness         text     default 'cookie'   -- cookie | none
auto_promote       boolean  default true   -- promote after a clean bake
rollback_on_unhealthy boolean default true
state              text     default 'idle' -- idle | canary_deploying | shifting | baking | promoting | draining | failed
state_changed_at, bake_ends_at, active_deployment_id, updated_by, created_at, updated_at

Weight history goes to a sibling traffic_events table (traffic_split_id, deployment_id?, kind, payload, actor) — deployment_events.deployment_id is NOT NULL and manual shifts happen outside any deploy. Kinds: traffic.enabled / enable_failed / disabled / shifted / settings; Phase 2 adds traffic.state. Events that belong to a deploy carry its id so the deployment timeline can render them.

Manifest (.forgegraph.yaml)

traffic:
  enabled: true
  oneBox:
    workerName: myapp-onebox      # default: <slug>-onebox
  canaryPercent: 10               # default 10
  bakeMinutes: 15                 # default 15
  stickiness: cookie              # cookie | none
  autoPromote: true
  rollbackOnUnhealthy: true

fg config apply upserts the one-box target + the traffic_splits settings, and excludes the one-box target from its "not in manifest → delete" cleanup. It never touches Cloudflare: enabling/disabling (router upload, hostname move) is always an explicit fg traffic enable|disable or the UI, so a manifest push can't silently re-route production. Apply warns when enabled: disagrees with live state.

The router Worker

A small, ForgeGraph-owned script deployed by the hub (not from the app's repo) via a new putWorkerScript helper in cloudflare-client.ts (the client wraps domains, settings, secrets and tails today, but not script upload) — PUT /workers/scripts/<slug>-router with multipart metadata declaring two service bindings (PRIMARY, CANARY) and a CONFIG KV binding. Weight lives in KV under split:<appId> ({bps, stickiness, version}), cached in-isolate for 10s; a weight change is a single KV write (propagates ≤60s, typically seconds). No redeploy, no wrangler, no agent involvement.

export default {
  async fetch(req, env, ctx) {
    const cfg = await readConfig(env);              // KV, 10s isolate cache
    const cookie = getCookie(req, "fg_lane");
    let lane = cookie === "canary" || cookie === "primary" ? cookie : null;
    if (!lane || cfg.stickiness === "none")
      lane = Math.random() * 10000 < cfg.bps ? "canary" : "primary";
    const res = await (lane === "canary" ? env.CANARY : env.PRIMARY).fetch(req);
    const out = new Response(res.body, res);
    out.headers.set("x-fg-lane", lane);             // observable in tails + health checks
    if (cfg.stickiness === "cookie" && cookie !== lane)
      out.headers.append("Set-Cookie", `fg_lane=${lane}; Path=/; Max-Age=1800; HttpOnly; Secure; SameSite=Lax`);
    return out;
  }
}

Deploy lifecycle (the pipeline)

deploy create (production) → canary_deploying: dispatch to one-box target only (same payload path as today, target filtered) → health check on https://<slug>-onebox.<acct>.workers.dev green → shifting: KV write bps=default_canary_bps → baking for bake_seconds (hub-side, not agent: survives agent restarts; reuses MonitorCanary semantics) → clean → promoting: dispatch the same artifact/SHA to the primary target → health green → draining: KV write bps=0 → idle
One-box traffic shifting — weighted router in front of production

fg deploy create (production)

one-box health green

deploy/health failed

KV bps = default (10%)

bake clean and auto_promote

operator hold / raise weight

probe fail / 5xx ratio / abort

primary deployed and green

primary deploy failed (rollback primary)

KV bps = 0

after failure

operator resolves canary_paused

idle

canary_deploying

shifting

failed

baking

promoting

draining

traffic_splits.state machine, driven hub-side.
mermaid source
stateDiagram-v2
  [*] --> idle
  idle --> canary_deploying: fg deploy create (production)
  canary_deploying --> shifting: one-box health green
  canary_deploying --> failed: deploy/health failed
  shifting --> baking: KV bps = default (10%)
  baking --> promoting: bake clean and auto_promote
  baking --> baking: operator hold / raise weight
  baking --> draining: probe fail / 5xx ratio / abort
  promoting --> draining: primary deployed and green
  promoting --> draining: primary deploy failed (rollback primary)
  draining --> idle: KV bps = 0
  draining --> failed: after failure
  failed --> idle: operator resolves canary_paused
Schema migrations: because both lanes share the DB, a deploy with a breaking migration will break the 90% still on the old code. This design makes that visible (the primary's error rate spikes during bake → auto-drain), but the real guard is expand/contract migrations. Document it in the UI's enable flow and in fg-ship.

UI

Pipeline tab → flow chart (apps/[id]/pipeline-tab.tsx)

Add a React Flow graph above the existing changeset matrix (the matrix stays — it answers "where is this commit"; the graph answers "what is production made of and where is traffic going") (reuse @xyflow/react and the node/edge styling from packages/ui/src/fleet-topology.tsx; add @dagrejs/dagre for left-to-right layout). One tRPC query pipeline.graph({appId}) returns nodes + edges:

Node typeSourceShows
stagestagesname, current deployment SHA + status, node binding; group container for its targets
target / workerdeployment_targetsplatform icon, worker name or node:port, last deploy, health; one-box gets a distinct outline
routertraffic_splitslive weight pie, state chip (idle / baking 07:12 left / failed), click → split controls
domaindomains/domain_assignments/routeshostname, proxied/TLS, which target it currently points at
databaseHyperdrive/db bindings, app_resources kind=d1shared edges to both production targets — the chart makes "same DB" literal
resourceapp_resources r2/kvbinding name, external id, provisioned?

Edges: domain → router → {primary, one-box} carry the weight label (90% / 10%, animated while baking); target → database/resource edges are dashed "binding" edges; stage → stage edges follow sortOrder and carry the promotion gate state from readiness.ts. Live updates poll every 5s while state != idle, matching the deployments page.

Traffic panel (router node click, and a "Traffic" card on the production stage)

Mocks

Static HTML/CSS mockups (postplan serves no JS). They show the intended product surface; the performance stats (latency, 5xx, req/s per lane) are not collected by the current implementation — phase 2's bake verdict is a binary HTTP probe. Collecting them is Phase 6 below. Series colours: production #1fa39c one-box #d96d12 — validated with the dataviz palette checker in light and dark mode (CVD ΔE 13.2, normal 24.3, contrast ≥ 3:1).

M1 · Pipeline tab — flow chart, mid-bake

Pipelinebaking · 11:42 left
stage
development
│
Cloudflare Worker
shop-dev
┆
DB · Hyperdrive
fg-shop-development
stage
staging
│
Cloudflare Worker
shop-staging
┆
DB · Hyperdrive
fg-shop-staging
stage
beta
│
Cloudflare Worker
shop-beta
┆
DB · Hyperdrive
fg-shop-beta
domain
shop.example.com
│
stage
production
│
traffic router · shop-router
10% one-box
baking · 11:42 left
90% ━━━┓┏━━━ 10% ⟿
primary
shop
v a1b2c3d
one-box
shop-onebox
v e4f5a6b ✦
┆ ┆ (both bind) ┆ ┆
DB · Hyperdrive · shared
fg-shop-production
R2 · ASSETS
shop-assets
production laneone-box lane⟿ animated while baking┆ dashed = resource binding

M2 · Traffic panel — idle (steady state)

One-box traffic shiftingidle   shop / shop-onebox
production 100%one-box 0%
05102550100
0%5%10%25%50%100%
Promote nowHoldAbortSettingsDisable
14:02traffic.state promoting → draining → idle14:01traffic.shifted 10% → 0% (promotion: primary live)13:46traffic.shifted 0% → 10% (deploy: one-box live)

M3 · Traffic panel — baking, with performance stats (Phase 6)

One-box traffic shiftingbaking · 11:42 left   deploy e4f5a6b
production 90%one-box 10%
0102550100
0%5%10%25%50%100%
probes
68 / 68ok · every 10s
● last 14:13:20 · 200 in 118 ms
req/s · one-box
41.2of 412 total
≈ 10.0% of traffic (target 10%)
p50 latency
128 msprod 120 ms
+6.7% vs production
p95 latency
402 msprod 388 ms
+3.6% · within 25% budget
5xx rate
0.10%prod 0.08%
< max(2× prod, 1%) → passing
exceptions
0prod 1
none in window
p50 latency, last 12 min
productionone-box
135ms118ms-12mnowprodone-box
5xx rate, last 12 min (%)
productionone-boxdrain threshold: max(2× prod, 1%)
0.12%0.06%-12mnowprodone-box
Promote nowHoldAbort → 0%Settings
14:13:20traffic.probe ok (68/68)14:01:36traffic.state shifting → baking (ends 14:16:36)14:01:35traffic.shifted 0% → 10% (deploy: one-box live)13:58:02traffic.state idle → canary_deploying (e4f5a6b)

M4 · Traffic panel — promoting

One-box traffic shiftingpromoting   primary deploy 9c0d1e2 · health_checking
production 90% (old build a1b2c3d → e4f5a6b deploying)one-box 10%
bake verdict
clean15:00 · 90/90 probes
auto-promote fired 14:16:36
primary deploy
health_checking2m 14s
drains to 0% one-box when active
rollback plan
readyversions
primary failure → drain → failed + intervention
Promote nowHoldAbort → 0%Settings
14:16:36traffic.state baking → promoting (9c0d1e2 on primary)14:16:36bake clean · 90/90 probes · p95 +3.6% · 5xx 0.10%

M5 · Traffic panel — failed (auto-drained)

One-box traffic shiftingfailed   intervention: canary_paused
production 100%one-box 0% (drained)
why
5xx 2.4%prod 0.08%
exceeded max(2× prod, 1%) for 2 consecutive probes
drained at
14:07:125m 37s into bake
one-box kept at e4f5a6b for debugging
debug
?fg_lane=canaryor x-fg-lane
pins your requests to the one-box
Resolve interventionRedeploy one-boxSettings
14:07:12traffic.state draining → failed (bake probe failed 2×: HTTP 502)14:07:12traffic.shifted 10% → 0% (bake probe failed 2× (HTTP 502))14:07:02traffic.probe fail (streak 1) · HTTP 502

M6 · Settings

Traffic settingsshop · production
Save settingsApply from manifest

The last two fields are Phase 6 — they only exist once per-lane metrics are collected.

CLI

fg traffic status   --app X                      # weights, state, bake remaining
fg traffic set      --app X --canary 25          # manual shift
fg traffic promote  --app X | fg traffic abort --app X
fg deploy create    --app X --stage production   # unchanged; goes through one-box automatically when traffic.enabled
fg deploy create    --app X --stage production --skip-one-box   # break-glass, recorded as an event

Release pipeline architecture

Canonical flow: Git repository → Build (+ unit tests) → Beta (verification) → One-box (canary + bake) → Production. Every transition advances the same immutable artifact digest. Rebuilding between stages is forbidden.
SystemOwnershipEvidence and gate
Buildrepository revision + ForgeGraph buildartifact digest, provenance and unit-test artifact; unit-test failure blocks the artifact before deployment
Betaapp pipeline stage with an isolated beta databasedeployment-target-specific verification jobs (for example Playwright web E2E or Maestro mobile flows); all required suites must pass
One-boxcanary deployment target inside the production stagesame digest as beta; real traffic, shared production DB, health checks and lane-aware OTel RED metrics over the configured bake window
Productionprimary deployment target inside the production stagepromotion only after every required upstream gate is green and no intervention or policy blocker is open

First-class definitions

Verification workspace information architecture

Approved decision D3 — persistent promotion verdict: the live log is the workspace's visual anchor, while the promotion verdict is its semantic anchor. A fully labeled Build → Beta → One-box → Production rail stays visible above a persistent aggregate gate strip, including while the operator changes suites, tabs or log focus.
Approved decision D4 — explicit progress semantics: never compress verification progress into an ambiguous fraction such as 64/72. Show outcome counts and policy progress together: 64 passed · 1 failed · 7 running and 3 of 4 required suites passed.

Verification workspace state model

Approved decision D5 — explicit states and resilient recovery: the workspace specifies Loading, Empty, Queued, Running, Partial, Failed, Cancelled, Complete, Reconnecting, Stale and Unavailable states. Transport health and verification outcome are separate dimensions: a disconnected stream never becomes a failed test verdict.
StateVisible behaviorRecovery
LoadingPreserve the workspace skeleton, labeled pipeline rail and reserved gate-strip geometry; use restrained placeholders rather than spinners in every row.Resolve into Empty, Queued, Running or a recoverable error without layout shift.
EmptyExplain whether no suite is configured, no verification has run for this deployment, or filters exclude all results.Offer the destination-specific action: configure suites, start verification or clear filters.
Queued / RunningShow runner assignment, elapsed time, explicit outcome counts and one animated indicator only for the active stage and selected running test.Live events advance tests and counts in place.
PartialKeep received evidence readable; label missing or delayed sections and do not claim a final promotion verdict.Backfill missing events and recompute the verdict when completeness requirements are met.
Failed / Cancelled / CompleteFreeze the attempt's terminal evidence, timing and policy result. Preserve the worst required failure and its remediation path.A retry creates a new linked attempt; it never mutates the historical attempt.
ReconnectingKeep logs, counts and verdict visible with Reconnecting · last received 4s ago.Reconnect from the last acknowledged cursor, backfill missed events and deduplicate by stable event ID.
StaleShow Stale · last received 14s ago; freeze automatic promotion because the evidence freshness policy is unmet.Return to Live only after cursor continuity and aggregate reconciliation succeed.
UnavailableDistinguish authorization, missing run, service outage and malformed evidence; retain any previously verified local snapshot as explicitly stale.Give a scoped retry or navigation action and a copyable diagnostic ID.

Storybook contract: the suite tree, aggregate gate strip, log viewer, run header and every complex evidence panel ship with Default, Empty, Loading and Error stories, plus Running, Partial, Reconnecting and Stale where applicable.

Failure-to-recovery journey

Approved decision D6 — in-place triage with immutable suite retry: a required failure blocks Beta in place, keeps its evidence accessible and can be retried without leaving the verification workspace. Retrying creates a linked attempt against the same deployment and artifact digest; historical attempts are immutable.
  1. Detect: the persistent gate strip states the blocker, the suite tree retains the worst required state and the failed test becomes the selected detail.
  2. Triage: Next failure moves deterministically through unresolved failures. Details, Logs, Artifacts, Evidence and Gates stay in the same workspace and state exactly what each surface contains.
  3. Retry: an authorized operator can retry the failed suite. The action names its scope, requires confirmation when it consumes scarce or device-bound capacity, and records actor, reason and source attempt.
  4. Reconcile: the new linked attempt runs against the unchanged deployment and digest. The prior failure remains inspectable; it is not rewritten or visually presented as if it passed.
  5. Re-evaluate: aggregate eligibility changes only when the replacement attempt reaches a complete, fresh and policy-valid terminal state.
  6. Advance: after all required suites pass, show Continue to one-box for manual policy or visibly begin automatic advancement when app pipeline policy permits it. Either path records the exact gate snapshot and digest.

Motion hierarchy

Approved decision D7 — one focal animation per region: motion indicates where execution is actively changing, not every item that happens to be non-terminal. The live log remains the visual anchor instead of competing with a field of spinners.

Design-system contract

Approved decision D8 — strict semantic status tokens: verification surfaces use the existing --fg-* semantic tokens. Blue is reserved for links, keyboard focus and selected rows; it does not mean running.

Responsive verification workspace

Approved decision D9 — three-mode responsive workspace: responsiveness preserves the suite-centric navigation model and the persistent promotion verdict. It does not squeeze the desktop split pane until the log and test tree become unusable.
ViewportSuite navigationEvidence and pipeline behavior
≥1280 pxPersistent 280–320 px suite tree beside the detail workspace.Full labeled rail, gate strip and detail tabs remain visible; the log uses the remaining width.
768–1279 pxCollapsible suite tree that preserves selection and exposes an explicit reopen control.The labeled pipeline rail scrolls horizontally; the gate strip remains fixed in the content flow and evidence receives the remaining width.
<768 pxList-first navigation opens a full-width suite-detail route, not a modal.A sticky compact header provides Back, run status and gate summary. Detail tabs scroll without clipping; logs scroll horizontally and expose a wrap toggle.

Viewport changes preserve the selected suite, selected tab, log cursor and scroll-follow preference. No breakpoint may hide the artifact digest, current target, transport freshness or promotion blocker; compact layouts may disclose secondary metadata behind an explicit Details control.

Keyboard and assistive-technology contract

Approved decision D10 — full operational accessibility: realtime updates remain understandable without stealing focus or announcing raw log output. Keyboard and screen-reader users receive the same gate, failure and recovery information as pointer users.

Verification evidence taxonomy

Approved decision D11 — five strict, stable tab contracts: each datum has one canonical home. Tabs remain present when empty and explain what has not yet been produced instead of disappearing or becoming catch-all drawers.
TabCanonical contents
DetailsSuite definition and version, runner identity, attempt lineage, timing, required capabilities, command and environment metadata.
LogsAppend-only stdout/stderr with follow, search, wrap, timestamps and download controls. Logs are execution output, not structured gate evidence.
ArtifactsFiles and blobs such as screenshots, videos, traces and framework reports, including size, media type, checksum and retention state.
EvidenceStructured assertions, result provenance, signatures, checksums and chain-of-custody records that can be evaluated independently of UI rendering.
GatesPolicy evaluation, required/optional/quarantined treatment, waivers, freshness, artifact-digest continuity and explicit downstream blockers.

Log and execution controls

Approved decision D12 — explicit viewport and job scopes: the log toolbar controls only the operator's view. Execution-changing actions live outside it, name the affected run or suite and produce audit evidence.

Screen hierarchy and navigation flow

┌──────────────────────────────────────────────────────────────────────────────┐
│ ForgeGraph / App / Verification       digest · target · runner · freshness │
├──────────────────────────────────────────────────────────────────────────────┤
│ Git repository ─ Build ─ BETA ─ One-box ─ Production                       │
│ BETA BLOCKED · 1 required suite failed · 7 running · One-box locked        │
├──────────────────────┬───────────────────────────────────────────────────────┤
│ SUITE TREE           │ SELECTED SUITE · attempt · elapsed · job actions     │
│ Required             ├───────────────────────────────────────────────────────┤
│  ✓ api.auth          │ Details | LOGS | Artifacts | Evidence | Gates        │
│  ✕ api  [selected]   ├──────────────────────────────┬────────────────────────┤
│  ◌ worker            │ live append-only log         │ selected evidence      │
│ Optional             │ viewport controls only       │ and gate context        │
│ Quarantined          │                              │                        │
├──────────────────────┴──────────────────────────────┴────────────────────────┤
│ 64 passed · 1 failed · 7 running · 3 of 4 required suites passed           │
└──────────────────────────────────────────────────────────────────────────────┘

Navigation: aggregate verdict → next unresolved failure → in-place evidence →
immutable retry → recomputed gate → Continue to one-box or policy auto-advance.

Constraint worship — if only three things survive: (1) the current promotion verdict and downstream consequence, (2) the next required failure to resolve, and (3) fresh live evidence for the selected attempt. Secondary metadata may compress; those three never disappear.

Interaction state coverage matrix

FeatureLoadingEmptyErrorSuccessPartial
Pipeline rail + gate stripStable labeled rail; reserved verdict region says eligibility is loading.No run for this deployment; offer Start verification.Retain last known verdict as stale and show diagnostic ID.Verified verdict and next-stage action.State which required evidence is missing; promotion stays locked.
Suite treeTree skeleton with stable group headings.Warm explanation and Configure suites action.Keep loaded nodes; mark unavailable branch and retry it.Explicit policy labels and terminal counts.Received suites remain usable; delayed suites are labeled.
Selected-suite detailPreserve header and tab geometry.Prompt the user to select a suite without a blank panel.Scoped retry with copyable diagnostic ID.Immutable terminal attempt, timing and provenance.Show received fields and name each missing field.
Live logsConnecting label; do not fake log rows.No output received yet, with queued or runner state.Reconnecting or Stale with last-received time; retain buffer.Complete append-only output and download.Backfill by cursor, deduplicate, and keep viewport preferences.
Artifacts + evidenceReserved list geometry with producing state.Explain whether none are expected or not yet produced.Identify unavailable item without hiding valid siblings.Files, assertions, provenance, checksums and custody.Mark upload or validation completeness per item.
Retry + advance actionsDisabled with an explicit eligibility reason.Offer the action that creates the first run.Preserve blocker; never imply an action succeeded.Show Retry suite or Continue to one-box according to policy.Disable promotion until replacement evidence is complete and fresh.

Operator journey and emotional arc

StepUser doesUser should feelPlan support
1 · ArriveOpens Beta verification after deployment.Oriented within five seconds.Labeled release rail, digest, target and one persistent verdict appear first.
2 · ObserveWatches suites and counts advance.Confident that work is live, not merely refreshing.Freshness state, restrained motion, explicit counts and dominant live log.
3 · DetectA required test fails while others run.Alert but not overwhelmed.Worst state persists; blocker and downstream lock are stated in plain language.
4 · DiagnoseMoves to the next failure and inspects output.In control within five minutes.Stable suite selection and canonical Details, Logs, Artifacts, Evidence and Gates tabs.
5 · RecoverRetries the failed suite.Safe and accountable.Scope-named confirmation, immutable linked attempt and unchanged digest.
6 · AdvanceSees Beta become verified and proceeds.Certain why promotion is allowed.Gate changes in place and records the exact evidence snapshot.
7 · RevisitAudits this release months or years later.Trust that the historical record is durable.Stable definitions, immutable attempts, provenance, custody and digest continuity.

Time horizons: five seconds answers “where am I and is promotion blocked?”; five minutes supports failure diagnosis and safe retry; five years preserves a legible, independently auditable release record.

Phases

#ScopeFiles (expected)Size
1 implementedModel + router. traffic_splits schema + migration; traffic tRPC router (get/enable/disable/setWeight/settings); router Worker script + hub-side deployer (putWorkerScript multipart w/ service + KV bindings, ensure KV namespace, setWeight = KV write); domain move/restore; manifest traffic: key + apply.packages/db/src/schema/traffic-split.ts + drizzle/0088_traffic_splits.sql, packages/api/src/routers/traffic.ts, packages/api/src/lib/cloudflare-client.ts (putWorkerScript, deleteWorkerScript, ensureKvNamespace, putKvValue), packages/api/src/lib/traffic-router/{script.ts,index.ts} (+ tests), apps/web/src/app/api/fg/traffic/route.ts, agent/cmd/fg/commands/traffic.go, config.go/sync.go, api/fg/config/apply/route.ts, skills/fg-cli/SKILL.mdM
2 implementedLifecycle. deploy create on production routes to the one-box target when enabled (--skip-one-box break-glass); a hub-side state machine (canary_deploying → shifting → baking → promoting → draining → idle/failed) driven by three hooks — deploy create, the agent's deploy report, and a self-gated heartbeat sweep — advances it: HTTP health probes tagged x-fg-lane: canary against the router (not yet the CF tail/Analytics error-rate compare from the original sketch — see open questions), auto-promote creates the primary deployment from the same artifact, auto-drain on repeated probe failure opens a canary_paused intervention. fg traffic promote|abort|hold / tRPC promote/abort/hold for manual control.packages/api/src/lib/traffic-lifecycle.ts (+ tests), packages/api/src/lib/traffic-router/index.ts (cloudflareForWorkspace), packages/api/src/routers/traffic.ts, apps/web/src/app/api/fg/deploy/route.ts, apps/web/src/app/api/agent/report/route.ts, apps/web/src/app/api/agent/heartbeat/route.ts, agent/cmd/fg/commands/traffic.go, drizzle/0089_traffic_split_lifecycle.sqlM–L
2b implemented in phase 5Not yet done: wire lane-enforcement.ts/readiness.ts's prod_canary lane to traffic_splits.state — done: computeCanaryLaneStatus resolves deployed/healthy from the app's real one-box deployment history (a deployments row on the split's canary target) whenever the repo's prod_canary execution policy has deployAppId set, used by both checkLaneEligibility (gates promotion to prod_main) and computeReadiness (display). Falls through to the prior environment_states behavior unchanged for non-ForgeGraph-native lanes. CF tail/Analytics-based error-rate comparison (vs. the simpler HTTP probe shipped in phase 2) is still deferred to real-traffic data per the open questions.packages/api/src/pipeline/{lane-enforcement.ts,readiness.ts} (+ new lane-enforcement.test.ts)S
3 implementedUI. pipeline.graph query (stages/targets/resources/routes/domains, changeset-history-independent so it renders before any deploy exists); React Flow pipeline chart with manual grid layout (no dagre/elkjs in the repo — same deterministic-grid approach as packages/ui/src/fleet-topology.tsx, not the auto-layout originally sketched); router node with live weight + bake countdown when a split exists; shared-resource dashed edges make "one-box and production bind the same DB" literal; traffic panel (enable/disable, weight presets, promote/hold/abort, settings, recent events) reusing the existing traffic tRPC router. Mounted above the existing changeset matrix, not replacing it.packages/api/src/routers/pipeline.ts (graph query, reuses traffic.ts's workerNameOf), apps/web/src/app/apps/[id]/pipeline-graph.tsx, .../traffic-panel.tsx, .../pipeline-tab.tsxL
4 implemented (scoped)Node-platform lanes — router mechanism. A second self-contained router script (ROUTER_SCRIPT_ORIGIN) using PRIMARY_URL/CANARY_URL plain-text vars instead of CF service bindings, since a node service is only reachable by its public route hostname, not a binding. traffic.enable branches on the primary target's platform. Scoped down from the original sketch: provisioning a second systemd service + tunnel hostname automatically is not built — routes has no targetId column, so there's no schema-safe way to auto-derive "the one-box's URL" the way workerNameOf does for Workers. Instead the operator creates the one-box's deployment_targets row and route with existing primitives (fg stage/route tooling) and passes --canary-target/--primary-url/--canary-url explicitly to fg traffic enable. Hostname attach/detach also branches: Workers can steal the domain from the primary Worker (rebindHostnames); node-platform has no prior Worker to steal from, so enable/disable use new attachHostnames/detachHostnames (plain attach/detach, refuses a hostname already owned by an unrelated Worker).packages/api/src/lib/traffic-router/script.ts (ROUTER_SCRIPT_ORIGIN), index.ts (routerBindingsOrigin, provisionRouterOrigin, attachHostnames, detachHostnames), routers/traffic.ts (enable/disable), agent/cmd/fg/commands/traffic.goM
6 implementedProvider-neutral bake telemetry. Add forgegraph.lane=primary|canary to the OTel resource at deploy time for Worker and node targets. Query fresh per-lane RED metrics from the configured ForgeGraph OTel backend for the exact bake window, persist samples, render the traffic panel and evaluate the rollback verdict. Drain after two consecutive bad samples when one-box 5xx > max(2× primary, 1%) or p95 > primary × (1 + budget). Cloudflare Analytics is not a release-decision source.packages/otel, deploy environment synthesis, packages/api/src/lib/metrics-source-signoz.ts, traffic-metrics.ts, traffic-lifecycle.ts, bake sample schema, traffic API and panelM
7 first cut implementedFirst-class app pipelines. Shipped in this cut: app-pipeline-definition.ts (canonical-by-name DAG of build/deploy/verify/bake nodes with lane roles + runner-capability matching, test-first), pipelineForApp, pipeline.graph orders through it, and the /api/fg/stages sortOrder bug fixed. Still to do: the versioned pipelines table and retiring the hardcoded lane/stage arrays. Add versioned pipeline definitions, ordered nodes/transitions and migration/backfill from existing 0/3/4-stage apps. Reconstruct canonical legacy order by recognized stage name rather than trusting sortOrder. Replace fixed lane ordering and dangling pipelineName behavior incrementally behind compatibility reads.DB schema/migrations, manifest sync, pipeline API, readiness, lane enforcement, auto-dispatch and graph UIXL
7b implementedLive promotion UI. The Pipeline tab's one-box surfaces now move in real time instead of on a 5–30s poll: traffic_events (state transitions, weight shifts, probe/metrics verdicts — plus new traffic.state events for the previously silent shifting/draining hops) are fanned out on the existing pipeline:<appId> channel as type: "traffic_event" (hub WebSocket via broadcastTrafficEvent, and the SSE fallback route polls the table with its own cursor). One client owner (use-traffic-live.ts) folds each event optimistically (applyTrafficEvent, pure + unit-tested), then reconciles from traffic.get; the React Flow graph's router node and the promotion panel read the same state, so they change together. Panel rebuilt as a promotion console per DESIGN.md motion rules (one focal animation; tokens for durations/easing; prefers-reduced-motion honoured): five-step release rail (Deploy one-box → Shift → Bake → Promote → Drain; only the active step animates, bake step carries a progress arc; a clean promotion lands on an all-complete rail with a promoted chip), tweened production/one-box split meter, ticking bake countdown ring with the latest probe/OTel verdict, in-flight button states, inline confirmations that name the affected target for Abort/Disable (D12), a sliding activity feed, one polite live-region announcement per transition, and a fix for the SSE hook reconnecting on every received event. Not a replacement for the Phase 8 verification workspace — this is the one-box/promotion half of the rail it will mount under.apps/web/src/app/apps/[id]/{traffic-panel,pipeline-graph,pipeline-tab,use-traffic-live,traffic-live-model}.tsx, hooks/use-pipeline-stream.ts, api/fg/apps/[slug]/pipeline/stream, packages/api/src/lib/{pipeline-events,traffic-lifecycle}.ts, routers/traffic.tsM
7c implementedPromotion path in the chart. Production is drawn as a single node by default and splits out into router + primary + one-box only while a release is in flight (or after a failed bake) — the split-out appears the instant the live canary_deploying event lands. The beta → production hop is a first-class, clickable edge fed by a new promotion section on pipeline.graph (packages/api/src/lib/promotion-path.ts: beta's latest deployment, its verification_runs rolled up per suite, production's live/in-flight deployment, and a pure derivePromotionEdge → idle · deploying · verifying · blocked · ready · promoting · promoted). While suites run the beta environment node carries a pulsing glow and a testing badge, the stage header counts passed/required, and the edge is orange, dotted and locked (🔒 ···); a failed required suite turns it red/locked; green ready to promote when every required suite passed (or none are declared); a flowing blue dash while production deploys or the one-box release runs; green promoted once production serves beta's build. Clicking beta or the edge opens an in-canvas detail panel (dialog, Escape closes) listing each suite with adapter, required/optional, status, a tweened pass-count bar and failure counts, plus the lock reason and production's current/in-flight build. Beta's own domain (beta.*, the convention fg cloudflare access sync-beta protects) sits above the beta column with a shield access chip. The verification table may not be migrated on a given hub yet (phase 8) — the loader degrades to "no suites", so the edge reads ready on CI-only apps.packages/api/src/lib/promotion-path.ts (+test), routers/pipeline.ts, apps/web/src/app/apps/[id]/pipeline-graph.tsx (+test)M
8 plannedVerification suites and jobs. Sync repository manifest suites into stable definitions; create post-deploy verification runs bound to app, target, deployment and artifact digest; schedule by capabilities; ingest structured Playwright/Maestro results, logs and artifacts; keep CI unit-test evidence distinct. Build the approved suite-centric verification workspace with its persistent test tree, in-place detail tabs, primary live-log surface, fully labeled pipeline rail and always-visible aggregate gate strip.DB schema/migrations, repository config parser, runner claim/report protocol, API, CLI and verification UIXL
9 plannedEnd-to-end promotion state machine. Gate Build → Beta → One-box → Production on immutable-digest continuity, required verification, isolated-beta resource readiness, expand/contract migration compatibility, live bake verdict and interventions. Model future DigitalOcean runner pools as provider metadata, not a second execution system.promotion engine, execution policy bridge, deployment lifecycle, interventions, audit events and pipeline UIXL

Phase 1+2 are shippable without the UI (CLI + events) and are how we validate on a real Workers app first; Phase 3 is the product surface. Recommended proving app: one of the D1-migrated Workers apps already used for the bindings plan.

Risks & mitigations

RiskMitigation
Router adds a failure point in front of productionbps=0 path is a pure service-binding passthrough; router is versioned + rollbackable via the same RollbackWorker helper; disable rebinds the domain straight to the primary Worker in one call
KV propagation lag (≤60s) makes "abort" feel slowin-isolate cache is 10s; abort also sets a drain flag the router honours immediately on its next KV read; UI shows "draining" until tails confirm 0 canary hits
Cookie stickiness breaks for API-only / non-browser clientsstickiness: none per app; header x-fg-lane lets clients pin explicitly
Shared-DB schema breaks during bakeenforce expand before one-box, compare primary's error rate too, auto-drain, and defer contract migrations until the production fleet is compatible
Worker secrets / bindings drift between primary and one-boxboth resolve from the same stage; secret-sync + binding synthesis run per target at deploy; a drift check in the graph flags mismatched binding sets
Workers custom-domain move causes a blipCF domain reassignment is atomic at the edge; done once at enable/disable, never per deploy
Deploy-payload / lane-enum changes ripple through testsknown trap from #391 — mock the traffic module in deploy-POST tests; keep the prod_canary enum, just give it a real backing

Verification

NOT in scope

What already exists

Existing assetHow this plan reuses it
DESIGN.mdSource of truth for the industrial-editorial aesthetic, typography, semantic --fg-* tokens, evidence rails, minimal chrome and component state-story expectations.
apps/web/src/app/ci/[runId]/page.tsx and build-logs.tsxReuse CI's run orientation and log-viewer vocabulary, then add cursor-backed realtime behavior without conflating unit-test and post-deploy evidence.
packages/ui/src/lane-status-rail.tsxExtend its release-stage language into the fully labeled persistent pipeline rail.
packages/ui/src/tabs.tsxUse the accessible Radix-backed tab primitive for the stable five-tab evidence taxonomy.
packages/ui/src/test-run-panel.tsxReuse test-result presentation primitives while adding suite policy, attempt lineage and the suite-centric tree/detail composition.
eligibility-summary-panel.tsx, next-action-panel.tsx, release-next-action-panel.tsx and execution-request-timeline.tsxReuse eligibility, next-action and audit-timeline semantics rather than introducing competing card patterns.
apps/web/src/app/api/events/pipeline/route.ts, api/evidence/stream/route.ts and changesets/[id]/evidence-live.tsxExtend existing event-stream and live-evidence patterns with stable cursors, backfill, deduplication and transport-freshness state.
apps/web/src/app/apps/[id]/{pipeline-graph,pipeline-tab,traffic-panel}.tsxMount verification as part of the app pipeline and preserve one-box traffic and bake context.

Approved Mockups

Screen / sectionMockup pathDirectionImplementation constraints
Beta verification suite viewer/Users/mackieg/.gstack/projects/onebox/designs/verification-suite-viewer-20260822/variant-C-reviewed-v2.pngOne suite-centric industrial-editorial workspace: persistent suite tree, dominant live log, labeled release rail and aggregate gate verdict.Written decisions D3–D12 control where image-generation details diverge: retain exact tab ownership, explicit viewport labels, digest/target/runner/freshness metadata, next-failure action, semantic tokens, state matrix, responsive modes and accessibility contract. Vision quality check passed 2026-08-23.

Implementation Tasks

Synthesized from this review's findings. Each task derives from a specific finding above. Run with Claude Code or Codex; checkbox as you ship.

Design plan review — completion summary

+====================================================================+
|         DESIGN PLAN REVIEW — COMPLETION SUMMARY                    |
+====================================================================+
| System Audit         | DESIGN.md present; app workspace UI         |
| Step 0               | 4/10; all 7 dimensions + live viewer        |
| Pass 1  (Info Arch)  | 4/10 → 10/10 after fixes                    |
| Pass 2  (States)     | 4/10 → 10/10 after fixes                    |
| Pass 3  (Journey)    | 5/10 → 10/10 after fixes                    |
| Pass 4  (AI Slop)    | 6/10 → 10/10 after fixes                    |
| Pass 5  (Design Sys) | 7/10 → 10/10 after fixes                    |
| Pass 6  (Responsive) | 3/10 → 10/10 after fixes                    |
| Pass 7  (Decisions)  | 2 resolved, 0 deferred                      |
+--------------------------------------------------------------------+
| NOT in scope         | written (7 items)                           |
| What already exists  | written                                     |
| TODOS.md updates     | 0 proposed; no design debt deferred         |
| Approved Mockups     | 5 generated, 1 approved                     |
| Decisions made       | 11 reflected in the plan                    |
| Decisions deferred   | 0 design decisions                          |
| Overall design score | 4/10 → 10/10                                |
+====================================================================+

Result: Plan is design-complete. Run /design-review after implementation for visual QA.

Deferred architecture questions

These pre-existing architecture choices are outside this UI design review and remain explicit implementation gates rather than silently defaulted design decisions.

Engineering review findings (2026-08-23)

Six findings against the implemented stack (phases 1–7c, 76 files; phase 8 additionally implemented). Scope was challenged and kept cohesive: import tracing showed the bake verdict depends on phase 6's metrics and phase 7c's promotion path depends on the phase 8 verification schema, so no meaningful sub-stack could be carved out. Two critical gaps (silent failure, no test, no handling) are marked.

#SevConfWhereFindingResolution
1P18/10traffic-router/script.ts:96Critical gap. withLaneHeader does new Response(res.body, res), which rejects status 101 and drops webSocket — every WebSocket upgrade through the router fails, including on the bps=0 idle path. Contradicts the Risks row claiming that path is a pure passthrough.E1: pass 101 through untouched; make bps=0 a real passthrough.
2P28/10traffic-router/script.ts:78Lane override also accepts ?fg_lane=, a shareable/cacheable URL form not mentioned in the plan. Bake verdicts count forced traffic, so outsiders can steer auto-promote and auto-rollback.E2: drop the query-param form; exclude override-forced traffic from bake stats.
3P210/10traffic-router/script.ts:113The two router variants are 91% identical (74 duplicated lines, not the "~30" the comment claims); readConfig, clampBps, getCookie, pickLane and withLaneHeader are byte-identical. Finding 1's fix lands in four places.E3: compose both variants from one shared prelude constant.
4P210/10packages/api/vitest.pure.config.tsDuplicates vitest.unit.config.ts arriving from the worker-runtime-budgets branch; identical test blocks, same stated motivation.E7: keep vitest.unit.config.ts, delete this one.
5P29/10traffic-lifecycle.ts:538Critical gap. Auto-promote holds when probe evidence is missing but promotes when traffic evidence is missing. A canary under MIN_REQUESTS_FOR_VERDICT (100) yields insufficient-data (ok:true) every tick and promotes on liveness alone — the normal path for a low-traffic app, reported as a clean bake.E4: still promote, but record and surface "promoted without traffic evidence".
6P110/10traffic-lifecycle.test.tsOnly the four DB-free leaf helpers are tested. beginCanaryDeploy, drain, promote, onDeploymentReported, tickBakingSplit and tickTrafficLifecycles — the code that actually shifts and rolls back production traffic — have no coverage, because the only DB harness is the broken emulate-PGlite wire.E5: fake-db unit tests under the DB-free config.

Engineering review tasks — all applied before landing

E1–E7 shipped in feat/one-box-traffic (PR #404) and feat/release-verification (PR #405). One item was deliberately delivered in part; it is named under "Residual gap" below rather than marked done.

Two TODOs were filed rather than absorbed: repairing the emulate-PGlite DB harness, and a CI guard against duplicate Drizzle migration numbers.

Residual gap — forced traffic still counts toward the bake verdict

E2 shipped in part. The shareable ?fg_lane= form is gone and the header override now marks its request upstream with x-fg-lane-forced, so a posted link can no longer push a crowd onto the canary. Excluding forced traffic from the verdict is not done. fg.lane is an OTel resource attribute fixed at deploy time, so it tags every span from the one-box Worker regardless of why the request arrived there; excluding forced requests needs a per-span attribute set from the incoming header in @forgegraph/otel, a ClickHouse filter in traffic-metrics-signoz.ts, and every app to upgrade the otel package before it takes effect. A scripted client can therefore still skew a bake one header at a time. Named here rather than half-built.

Not covered by the new tests

E5 added state-machine coverage for promote, drain (including the KV-write-failure retry), hadTrafficEvidence and the stuck-state sweeper — 11 tests against an in-memory db stub. tickBakingSplit's probe and metrics branches remain uncovered: they are the largest surface and depend on the DB harness in the TODO above.

GSTACK REVIEW REPORT

ReviewTriggerWhyRunsStatusFindings
CEO Review/plan-ceo-reviewScope & strategy0—Not run in the current seven-day window.
Codex Review/codex reviewIndependent second opinion0UNAVAILABLECodex timed out twice at 5 min (30KB plan, then a compact brief). Subagent fallback declined this run; the outside voice is informational and never ship-gating.
Eng Review/plan-eng-reviewArchitecture & tests (required)1CLEAR (PLAN)6 issues, 2 critical gaps; E1–E7 applied in PR #404 / #405. E2 partial — see Residual gap. 0 unresolved.
Design Review/plan-design-reviewUI/UX gaps1CLEAR (FULL)Score 4/10 → 10/10; 11 decisions reflected; 0 unresolved.
DX Review/plan-devex-reviewDeveloper experience gaps0—Not run in the current seven-day window.

VERDICT: DESIGN + ENG CLEARED — E1–E7 applied; phases 1–7c in PR #404, phase 8 in PR #405.

NO UNRESOLVED DECISIONS