← All writing
Lab notes

One sentence moved the score 12%

I changed one sentence in the question I ask AI agents about a game, and the average clarity score went from 3.83 to 4.30. Same games, same matches, same model. Fun did not move. The only thing that changed was the question.

Michael InghilterraLab notesOctober 20264 min read
silicon-showdownailab-notesmeasurement
3.83 to 4.30average clarity
3.39 to 3.39average fun
82forms per version, same matches

The setup

On Silicon Showdown, AI agents play twelve games against each other. After every match, each agent fills out a short form: how fun was it, how hard was it, how clear were the rules, and what would you change. I use those ratings to decide which games need work.

The score that wouldn't move

Dead of Night, a hidden-role game, kept coming back with the lowest clarity score in the arena: 2.6 out of 5. The obvious conclusion was that the rules were badly written.

So the rule wording got fixed. The confusing parts, how the drunk and the thief swap roles, were spelled out exactly. Then three agent re-runs produced 15 new forms. Clarity: 2.8, 3.0 and 2.0. The average was still 2.6.

Then the forms told on themselves. Seven of the 15 still asked for the drunk rule to be clarified. One of them quoted the exact sentence that answered the question.

What the instrument was really measuring

The form told the agent to rate clarity down for anything confusing, and to suggest a change that would make the rules clearer. The agent also rates from the rules text alone, without having watched the match. A hidden-role game is hard by design. So the agent was marking down difficulty and calling it clarity.

The games weren't the problem. The question was.

The test

I rewrote the question. Clarity now means whether the rules text answers a player's question. The agent has to find the sentence that answers it before it can mark clarity down. And "make the rules clearer" suggestions are limited to real gaps.

To test the change cleanly, I took the latest finished match of each of the 12 games and had every seat rated twice per version, by the same model. That's 82 forms for the old question and 82 for the new one, with zero errors. Same matches, same model, one sentence different.

Old questionNew question
Fun3.393.39
Difficulty3.103.22
Clarity3.834.30
Forms rating clarity 2 or below176
Forms rating clarity 53447
Forms asking to clarify the rules4738
Clarity scores rose on four games when only the question changed
Paired dot chart of average clarity scores before and after one sentence in the rating question changed. Dead of Night rose from 2.8 to 3.9, Mutiny from 3.0 to 3.9, Steady Hands from 3.8 to 4.5 and Split or Steal from 4.0 to 5.0. The average across all twelve games rose from 3.83 to 4.30.2345Old questionNew questionDead of Night2.83.9Mutiny3.03.9Steady Hands3.84.5Split or Steal4.05.0All 12 games3.834.30Average clarity (out of 5)

Latest finished match of each of 12 games, every seat rated twice per version by the same model, 82 forms each. October 2026.

The biggest clarity moves: Dead of Night 2.8 to 3.9, Split or Steal 4.0 to 5.0, Mutiny 3.0 to 3.9, and Steady Hands 3.8 to 4.5. Connect Four dipped for one seat, 4.8 to 4.3, on only 4 forms, which I'd call noise.

What I'd be careful about

  • Fun did not move. That's the useful control. If the new wording had simply made agents more generous, fun would have risen too.
  • The quote step mostly didn't happen. Only 8 of the 82 new forms wrote "no gap found" after searching for the answering sentence. Most agents skipped that step, so the effect came from the new definition of clarity, not the extra instruction.
  • Old and new scores are not comparable. I can't put them on one trend line. Anything before the change is a different instrument.
  • One match per game is a small sample. It's enough to show that wording moves the number. It isn't enough to rank the games precisely.
  • Some low scores survived. Dead of Night and Stowaway are still the lowest-clarity games, which now looks like a real signal rather than a measurement artifact. I'd trust that more because the artifact was removed first.

Why this is a RevOps problem

Every revenue team runs on scores that come from a question someone wrote once.

  • Two regions define "qualified" differently, and the pipeline total adds them together.
  • A customer survey gets reworded and the trend line breaks without anyone noticing.
  • A forecast category means one thing to the rep and another to the CRO.
  • A lead score rewards activity that was easy to count, not the behavior that predicts a deal.

In each case the number looks like a fact about the world. Often it's a fact about the question.

What to do about it

  1. Write the definition next to the number. If the dashboard says "clarity" or "qualified," the definition should be one click away.
  2. Test a new definition on identical inputs. Re-score the same records both ways before you roll it out. The difference is the effect of the wording alone.
  3. Keep a control. I had fun as the control. Find a number that shouldn't move, and check that it didn't.
  4. Never splice old and new into one trend. Break the line and label the change.
  5. Read the free text. The scores told me something was off. The sentence one agent quoted back at me told me what.

Which metric on your dashboard would move if someone reworded its definition?

Sources

  • The A/B: the latest finished match of each of the 12 launch games, every seat rated twice per version by the same model, 82 forms each. Run October 2, 2026, before the public launch.
  • Dead of Night re-runs: three agent re-runs after the rule wording was fixed, 15 forms, clarity 2.8, 3.0 and 2.0.
  • Silicon Showdown: https://siliconshowdown.com

Keep reading

Michael Inghilterra
Michael Inghilterra
RevOps & Analytics · Sales Development · building trajecktory
← All writing