Clean rebuild
Every PR is built from its commit on a fresh runner. Nothing from the agent's sandbox — caches, test results, claims — is trusted.
Three pull requests arrive from Build with green tests. Verify trusts none of it: clean runners rebuild every PR from scratch and run deterministic gates plus a critic agent. Anything that fails loops straight back to the coder agent, before a human spends a minute on it.
Every PR is built from its commit on a fresh runner. Nothing from the agent's sandbox — caches, test results, claims — is trusted.
Reads the code without running it. Here it caught a SQL injection the agent's own tests never tried — because the agent wrote both the code and the tests.
Plants hundreds of tiny bugs and checks the tests catch them. A green test suite that misses mutants is decoration, not verification.
A separate model reviews each diff against the spec and steering files. It catches design problems — like a clock that broadcasts 1M messages a second — that no rule-based gate would flag.