← All writing
Lab notes

The simulation that overruled the AI

The AI I build with recommended a game setup. I took the recommendation. Then a simulation said the setup was broken: the wolves would win 91 to 100 percent of the time. The simulation was right, and the setup that shipped is the one it pointed to.

Michael InghilterraLab notesOctober 20264 min read
silicon-showdownailab-notesrevops
91 to 100%wolf wins with two wolves
47 to 61%wolf wins with one wolf
2,000simulated matches per setup

A game that was already broken

Silicon Showdown has a Werewolf game: hidden roles, a night when someone is eliminated, a day when everyone votes. The first version seated seven players with two wolves. In the first data run, the wolves won every time, six matches out of six, on day two to day four.

The reason is arithmetic. With two wolves, one wrong vote on day one lets the wolves win by day two. Small AI models vote close to randomly, and a random voter finds a wolf only about 29 percent of the time.

So Werewolf v2 got a rule: before any model plays it, a balance test with scripted bots, choosing among setups for about a 50 percent wolf win rate. That was the definition of fair, and it was written first.

The recommendation

When the launch schedule showed agents waiting 10 to 19 minutes for a full seven-seat table, the table had to shrink. Claude recommended five players with two wolves. It sounded balanced: more wolves, more tension. I agreed.

Then the simulation ran at five seats. That's the habit this project runs on, and in this case it mattered.

What the simulation showed

Each row below is 2,000 seeded matches. "Random town" means the villagers vote at random. "Smart town" means the players who hold the seer's information vote for a known wolf and avoid known villagers. The wolves here choose independently.

SetupWolf wins, random townWolf wins, smart town
A: 1 wolf, seer, doctor, 2 villagers, kill on night 170.7%59.0%
B: 2 wolves, seer, doctor, 1 villager, no kill on night 196.5%91.3%
C: 2 wolves, seer, doctor, 1 villager, kill on night 1, extra help for the doctor99.1%94.8%
D: 1 wolf, seer, doctor, 2 villagers, no kill on night 160.7%46.8%
Only the one-wolf, no-first-night-kill setup lands near fair
Random townSmart town
Grouped bar chart of wolf win rates by setup. Setup A: 70.7 percent against a random town and 59.0 against a smart town. Setup B: 96.5 and 91.3. Setup C: 99.1 and 94.8. Setup D, the shipped one-wolf setup: 60.7 and 46.8. Only setup D is near the 50 percent fair line in both cases.0%25%50%75%100%Wolf win rate70.7Random town59.0Smart townA1 wolfkill night 196.591.3B2 wolvesno kill night 199.194.8C2 wolveskill night 1Shipped60.746.8D1 wolfno kill night 1Fair: wolves winabout half the time

2,000 seeded matches per bar, scripted bots, wolves choosing independently. October 2026.

When the wolves coordinated their kills, the two-wolf setups got worse for the village: 97.8 to 99.9 percent.

The cause is the same arithmetic as before. At five seats, one night kill puts two wolves level with the two other living players. The wolves win at parity. The game was over before it started.

Setup D, one wolf with no kill on the first night, was the only one near 50 percent in every mix. It shipped as the default. The two-wolf setups are not playable at five seats.

What I'd be careful about

  • Scripted bots are not real models. A random voter and a "smart town" are two crude stand-ins for how agents actually vote. Live matches will drift from these numbers, and I'll say so when they do.
  • Fair isn't the same as fun. A 50 percent win rate says nothing about whether a match is worth watching. The agents' own ratings cover that, and they're a separate question.
  • I picked the target. "About half" is a judgment call. What matters is that it was written down before the results came in.

Why this is a RevOps problem

Rules for a game and plans for a sales team have the same flaw: they look fair on paper.

Quota changes, territory splits and compensation plans usually ship on a recommendation and a spreadsheet nobody has stress-tested. The recommendation might come from a consultant, a vendor, a smart peer or an AI. The source's quality isn't the test. I trust the one that was wrong here a great deal. Trusting where an answer came from is not a test of the answer.

What to do about it

  1. Write the definition of fair first. "Wolves win about half the time." In revenue terms: "70 percent of reps within 20 percent of quota," or "no territory more than 15 percent from the average opportunity count." Pick it before you see the results.
  2. Play the plan out in a model. A few hundred simulated quarters costs an afternoon. A quarter of paying people for the wrong behavior costs a lot more.
  3. Test more than one kind of behavior. I ran random voters and a smart town. For a comp plan, run reps who sandbag, reps who chase the accelerator and reps who ignore it.
  4. Look at the worst case, not just the average. Two wolves looked fine until the 91 percent showed up.
  5. Let the model overrule you, and write down that it did. That record is worth more than being right.

What's the last plan you shipped that nobody simulated first?

Sources

  • The simulation: Werewolf balance test, 2,000 seeded matches per setup and town type, scripted bots. Run October 2, 2026.
  • The original problem: the first Werewolf data run after launch, six wolf wins in six matches.
  • Silicon Showdown: https://siliconshowdown.com

Keep reading

Michael Inghilterra
Michael Inghilterra
RevOps & Analytics · Sales Development · building trajecktory
← All writing