← The ADLC loop

Build: fast agents, safe boxes, tests first.

Plan v2 is approved. Three coder agents pick up the first wave of chess tasks — each in its own disposable sandbox — write the tests before the code, then prove the pieces work together in a full-stack preview environment.

    1. Agent sandboxes 1 per task · isolated · minutes
    2. Preview environment 1 per PR · full stack · hours
    3. CI gates Verify phase · independent re-run
    4. Staging → Prod Release phase · canary
    Tests appear here as the agents write them — before any implementation.
    Test pyramid · live
    0 tests
    Preview environment · pr-142.preview.internalNot started
      ♙ White · Chromium ♟ Black · WebKit
      E2E test waiting for players
      Move latency · 500 virtual playersp95 —
      150 ms
      050100150200 ms
      ← From Plan · v2 approved Click a test tab to read the code. Next phase Verify →

      Which test proves what

      TestWhat it provesIn the chess buildRuns in
      UnitOne function behaves, in milliseconds, with no network or database.The clock only ticks for the side to move (fake timers).Sandbox
      Property-basedA rule holds for thousands of generated inputs, not just the cases someone thought of.Random legal games never reach an illegal board.Sandbox
      IntegrationYour code works with real dependencies — not mocks.Matchmaking queue survives a restart, against a real Postgres container.Sandbox
      ContractServices agree on message shapes, so one side can't silently break the other.Every WebSocket move message matches the schema.Preview env
      End-to-endA real user journey works through real browsers.Two browsers play a game; Black sees White's move.Preview env
      LoadIt stays fast under realistic traffic.500 virtual players, move latency p95 < 150 ms (AC-1).Preview env
      Static + securityLint, types, SAST and dependency scans.Run by agents here, then re-run independently as gates.Sandbox → Verify

      Sandbox guardrails

      Isolated and disposable

      One container and one git worktree per task. When the PR is opened the sandbox is destroyed — nothing persists but the code and the test evidence.

      No secrets, narrow network

      No production credentials, ever. Outbound traffic is limited to an allowlist such as the package registry, so a confused agent can't reach anything that matters.

      Tests first, from the spec

      Agents write failing tests from the acceptance criteria before writing code. Green tests mean the spec is met — not just that the code compiles.