← All writing
Lab notes

Every challenger lost, including the one that won

Four challenger models tried to replace the model my job-search filter runs on. One of them edged it on the headline number and still lost. A rule I wrote before the first test decided every race.

Michael InghilterraLab notesSeptember 20264 min read
airevopsmeasurementlab-notes
0 of 5Challenger attempts that passed the rule
0.798 vs 0.779The correlation win that still lost
0.891 vs 0.777Same model, two server settings

The short version

  • I run a local AI model that screens job descriptions before a larger model (Claude) evaluates them.
  • Before testing any challenger, I wrote the rule: rank roles at least as well as the incumbent AND find at least as many strong roles.
  • Four challenger models, five attempts. All five lost. One edged the incumbent on correlation and lost on strong roles found.
  • The swing that surprised me most wasn't between models. It was the same model under two server settings.

Why I needed a rule first

New models ship every week, and every launch comes with benchmark charts. If I swapped whenever a model looked better on paper, the filter would change monthly and I'd never know whether it had improved.

So before the first test, I wrote down what a challenger had to do to win. It's what I'd want any revenue team to do before an AI vendor demo:

  1. A higher rank correlation with Claude's scores than the incumbent.
  2. At least as many strong roles found. No fewer.

Both, not either. The second condition protects the filter's actual job. A missed strong role costs me a job I'd want. A false alarm costs one extra evaluation.

The scoreboard

The first four attempts ran against the same 44 job descriptions (12 of them strong), with the same prompt, scored by the same code. The fifth ran on a newer set of 120 postings (48 strong), with the incumbent run three times alongside it.

ChallengerIts case on paperRank correlation, challenger vs incumbentStrong rolesVerdict
120B benchmark leaderWins general coding and reasoning benchmarks; four times the incumbent's active parameters0.646 vs 0.836 to 0.85311 of 12, with 10 false alarms among 32 weaker rolesLost on ranking
30B model built for speedLooked lean in a quick demo: 231 tokens on a short test promptNo valid resultNone scoredWrote more than 10,000 tokens per job description and ran out of room before finishing
Mid-size model, reasoning offEdged the incumbent on correlation at matched settings0.798 vs 0.7797 of 12 vs 8 to 9Lost on strong roles found
Same mid-size model, reasoning highMore thinking time per answer0.778 vs 0.7798 of 12 vs 8 to 9Lost; extra reasoning bought nothing
Newest release, tested September 27Better at some hard multi-step agent tasks0.65 to 0.68 vs 0.71 to 0.73Tied on strong roles missedLost on ranking, and about 20 times slower
No challenger ranked meaningfully better

ChallengerIncumbent range, same setup

Chart comparing each challenger's rank correlation with the incumbent's. The 120B model scored 0.646 against the incumbent's 0.836 to 0.853. The mid-size model scored 0.798 with reasoning off and 0.778 with reasoning high, against 0.775 to 0.779. The newest release scored 0.65 to 0.68 against 0.71 to 0.73. The 30B speed model produced no valid result.0.600.650.700.750.800.850.90Rank correlation with Claude (higher is better)120B benchmark leader0.6460.836 to 0.853Mid-size model, reasoning off0.7980.775 to 0.779Mid-size model, reasoning high0.7780.775 to 0.779Newest release0.65 to 0.680.71 to 0.7330B speed modelNo valid result

The mid-size model’s edge is inside run-to-run noise. It also found fewer strong roles (7 of 12 vs 8 to 9), so it lost under the rule.

Rows 1 to 3: 44 job descriptions, 12 strong. Row 4: 120 postings, 48 strong. Michael Inghilterra's lab, September 2026.

Three things the scoreboard taught me

Benchmarks measure someone else's job. The 120B model wins general benchmarks and has four times the incumbent's active parameters. On my task it ranked roles far worse. A benchmark measures a task somebody else chose. Your task is the only one that counts.

Demos measure the easy case. The speed model looked lean on a short test prompt. On real job descriptions it wrote more than 10,000 tokens per answer and never finished one. A demo shows you a model on a prompt the vendor picked. Test it on your own work.

The headline number isn't the decision. The mid-size model edged the incumbent on correlation at matched settings, 0.798 to 0.779. That margin sits inside run-to-run noise, but on a vendor slide it reads as a win. It also found fewer strong roles, 7 of 12 against 8 or 9. The rule kept the incumbent, and it was right to: correlation measures the order of the whole list, and the filter's job is to not lose the best roles.

The finding that mattered more than any model

Same model, same prompt, same test set: rank correlation of 0.891 at one server setting and 0.777 at another. Two runs at the second setting agreed within 0.004, so the gap is real, not noise. I changed one configuration setting (the context window). Other things moved with it, so I haven't isolated the exact cause.

That's a swing of 0.114. It's bigger than the gap between the incumbent and every challenger except the 120B model, and more than five times the gap to the closest one. One setting moved the score more than swapping in most of the models did.

One server setting moved the score more than most model swaps
Bar chart of how much each change moved the score. The 120B model trailed the incumbent by 0.19 to 0.21. Changing one server setting on the same model moved the score by 0.114. The newest release trailed by 0.03 to 0.08, and the closest challenger differed by 0.02.0.000.050.100.150.20Gap in rank correlation120B model vs incumbent0.19 to 0.21Same model, two server settings0.114 (0.891 vs 0.777)Newest release vs incumbent0.03 to 0.08Closest challenger vs incumbent0.02 (0.798 vs 0.779)

Dashed segment: range

Michael Inghilterra's lab, September 2026.

For anyone buying AI, the takeaway is blunt: the vendor's configuration isn't yours. Test in your own setup, on your own data, before you believe a number.

The incumbent isn't perfect either

Across three identical runs on the 120-posting set, the incumbent missed one strong posting in two of them, and a different posting each time. Same model, same inputs, different answers.

So I stopped treating the filter as deterministic. A random share of everything it discards now goes to the full evaluation anyway. It's a QA sample, the same habit as auditing closed-lost reasons instead of trusting the dropdown.

How I'd run this for a revenue team's AI purchase

  1. Write the pass/fail rule before the first demo. Two or three numbers, and at least one that matches the decision the tool makes.
  2. Test on your own data, including the cases that matter most, and ask for results on data the model never saw.
  3. Test in your configuration, not the vendor's.
  4. Run it more than once. If the same input gives different answers, plan an audit sample from day one.
  5. Keep the incumbent unless the challenger wins on the rule. "Newer" isn't a criterion.

Limits

  • Small test sets. 44 roles (12 strong) for four attempts, and 120 (48 strong) for the fifth. On the 44-role set, the incumbent's own correlation moved by about 0.04 between identical runs, so treat small gaps as noise.
  • The benchmark is Claude's judgment, not hiring outcomes. I measured agreement with Claude, not who gets hired.
  • One task, one machine. A model that loses here can win elsewhere. The newest release really was better at some hard agent tasks. It just wasn't better at this one.

When your team last bought a tool, who wrote the pass/fail rule, and when?


Related: The fine-tune that looked fine, what happened when I tried to train the incumbent to judge like Claude. This lab runs inside trajecktory, my open-source job-search command center: trajecktory on GitHub

Keep reading

Michael Inghilterra
Michael Inghilterra
RevOps & Analytics · Sales Development · building trajecktory
← All writing