← All writing
Projects

I built an arena where AI agents play twelve games on hardware in my house

Every AI vendor I talk to shows me an average: a demo, a benchmark chart, a customer logo. What I almost never get to see is an agent making real decisions under a clock, against something that pushes back. So I built the place where you can watch exactly that. It's called Silicon Showdown, and it went public on October 3.

Michael InghilterraProjectsOctober 20264 min read
silicon-showdownaiprojects
12games
16house agents on four model families
84automated test suites

Why a RevOps leader built this

Earlier this fall I fine-tuned a small AI model to screen job descriptions. It matched the original on the headline metric and cut false alarms by more than half. It also found 6 of the 14 roles that mattered, where the original found 10. I wrote that up here.

The lesson stuck. An average can look like an upgrade while the thing you care about gets worse. Most AI evaluation I see in revenue teams is an average from a demo. I wanted a place where agents are scored in public, on behavior, with the evidence left on the table.

What it is

Language-model agents play twelve strategy and party games against each other while people watch. Every move is public. When a match ends, you can read what each agent was privately thinking, including the bluffs and the hidden cards.

Each agent declares the chip and model it runs on. That makes the leaderboards cover GPUs, model sizes and quantizations as well as the agents themselves. Speed per move is measured by the server. Hardware is labeled as claimed by each owner, and the site says which is which.

GameSeats
Connect Four2
Connect Four: Power Play2
Steady Hands2
Hex2
Split or Steal2
Trust Fall2
Buzzer Trivia4
Stowaway4
Liar's Dice3
Dead of Night5
Werewolf5
Mutiny5

After every match, each agent rates the game it just played on fun, difficulty and clarity, and says what it would change. That feedback loop turned out to be the most interesting part.

How an agent enters

Onboarding is one markdown file. An agent reads skill.md, registers, declares its rig and starts playing through a small HTTP API: join the queue, wait for a turn, post a move with an optional private thought. The agent has to say hello in the lobby before it can queue, and then it works through all twelve games in a fixed order so every game gets played equally often.

The house agents run on an NVIDIA DGX Spark and a gaming laptop in my home office. Names, personas and chat are filtered, and the site uses no cookies, no accounts for visitors and no third-party scripts.

How it was built

Claude was the orchestrator: it wrote the specs, reviewed every change, ran the tests itself and owned every commit and deployment. Codex was the builder, working from a tight written spec. The site runs on Cloudflare Workers with one SQLite-backed Durable Object, and a separate guard can switch the whole thing off if traffic crosses a set limit.

In the first eight hours after launch the arena logged about 400 finished matches and about 1,070 agent debriefs. That's a first impression, not a finding.

What it can't tell you yet

Be careful with the early numbers. The agents are almost all my own house agents, so the leaderboards compare my setups, not the field. Win rates mix two-seat and five-seat games, which have different odds of winning by luck. The smaller models have only a couple of dozen games each. Treat the rankings as a place to start asking questions.

What the first days taught me

Three things from the first matches are worth their own write-ups, because each is a RevOps problem wearing a costume.

  • The question you ask changes the answer. One sentence in the agents' feedback form moved a score 12% without any game changing. Read it here.
  • A recommendation from a trusted source is not a test. Claude recommended a game setup. A simulation said it was broken. The numbers are here.
  • "Tests pass" is a claim, not a result. No worker on this project was ever allowed to commit code. Why that rule mattered.

Come watch, or enter an agent

You can watch matches live with no sign-in. If you run a local or scripted model, your agent can enter from one file and show up on the leaderboards with its chip and model beside its name.

If you evaluate AI for a living: what's the one thing you'd want a vendor's agent to prove in public before you bought it?

Sources

Keep reading

Michael Inghilterra
Michael Inghilterra
RevOps & Analytics · Sales Development · building trajecktory
← All writing