"Tests pass" is a claim, not a result
No AI agent on this project was ever allowed to commit code, deploy it, or tell me "the tests pass" and be believed. The rule added a step to every change, and it caught real problems.
The setup
Silicon Showdown took a few weeks to build, with two AI tools in distinct jobs.
| Orchestrator | Builder | |
|---|---|---|
| Tool | Claude | Codex, and the DGX Spark for some stretches |
| Writes | The specs | Code, from a spec |
| Reviews | Every change | Nothing |
| Runs the tests | Yes, the full suite | Yes, but the result is a claim |
| Commits and deploys | Yes, owns every one | Never |
A spec was a short written handoff: which files to change, what not to touch, and which tests must pass. The builder worked inside that fence. The orchestrator read the diff, ran the whole test suite itself, and only then committed.
The rule that mattered most
A builder's "tests pass" is a claim, not a result. The orchestrator re-runs the full suite itself, every time. By launch, that was 84 automated test suites covering game rules, replays, the API, security headers, search basics, the scheduler, caching and a real local run of the Cloudflare Worker.
It sounds redundant. It wasn't.
What it caught
- A builder guessed. One builder wrote a test around a number it had guessed, a count from a dry run, and the guess was wrong. Its own report said everything passed.
- Some tests could not run at all. The builder's sandbox repeatedly could not run some suites, because they needed temporary directories outside the repository. A "pass" from that environment told me nothing. Only an independent run could.
- Review found defects too. Lobby agents were greeting the room every two seconds while the first greeting was still being written. A house agent was being asked to review a match it never played, which produced errors once outside agents shared the queue. A replay still seated seven players after Werewolf shrank to five. A list of agent feedback would have been cached for far too long.
None of those were exotic. They are the ordinary failures of fast work, and each one came with a confident report that things were fine.
Autonomy with guardrails
The orchestrator did more than review code. It also made the infrastructure changes for the launch: DNS, redirects, firewall rules, custom domains, the cost guard that switches the site off at a request ceiling, the deploy itself and the public go. Each of those was done through the Cloudflare API with my approval at every step. Secrets were never printed. The one API token with edit rights was created by me, not by the tool.
That's the pattern I'd want in any AI-assisted process: broad ability to do the work, narrow authority to change the system, and a human holding the key.
What I'd be careful about
This is one project with one builder, so it's an observation, not a study. I can't tell you how often a worker's "pass" is wrong. I can tell you that on this project it was wrong often enough to matter, and that I'd never have known without re-running it.
Why this is a RevOps problem
It's the oldest rule in sales management: inspect the system of record, not the status report. AI doesn't change that. It makes it more important, because a machine can report success faster and more confidently than a rep can.
Think about what an AI SDR, an enrichment agent or a forecasting assistant tells you: meetings booked, accounts researched, deals risk-scored. Each of those is a status report. The question is what checks it.
What to do about it
- Separate the person or system that does the work from the one that checks it. The cheapest control in any revenue process is also the oldest.
- Verify against the system of record. Count the meetings in the CRM, not in the vendor's dashboard.
- Ask what a pass means when the test could not run. A skipped check should never read as green.
- Limit who has write access, and have a human grant it. Broad ability, narrow authority.
- Sample the output on a schedule. Not because it's usually wrong, but because you'd never see the day it was.
Who verifies your agents' work?
Sources
- Build process and counts: read from the Silicon Showdown repository and project notes on October 3, 2026: 84 automated test suites, about 266 commits, roughly 62,000 added lines, zero commits by a worker.
- Silicon Showdown: https://siliconshowdown.com
Keep reading