Every challenger lost, including the one that won
Four challenger models tried to replace the model my job-search filter runs on. One of them edged it on the headline number and still lost. A rule I wrote before the first test decided every race.
The short version
- I run a local AI model that screens job descriptions before a larger model (Claude) evaluates them.
- Before testing any challenger, I wrote the rule: rank roles at least as well as the incumbent AND find at least as many strong roles.
- Four challenger models, five attempts. All five lost. One edged the incumbent on correlation and lost on strong roles found.
- The swing that surprised me most wasn't between models. It was the same model under two server settings.
Why I needed a rule first
New models ship every week, and every launch comes with benchmark charts. If I swapped whenever a model looked better on paper, the filter would change monthly and I'd never know whether it had improved.
So before the first test, I wrote down what a challenger had to do to win. It's what I'd want any revenue team to do before an AI vendor demo:
- A higher rank correlation with Claude's scores than the incumbent.
- At least as many strong roles found. No fewer.
Both, not either. The second condition protects the filter's actual job. A missed strong role costs me a job I'd want. A false alarm costs one extra evaluation.
The scoreboard
The first four attempts ran against the same 44 job descriptions (12 of them strong), with the same prompt, scored by the same code. The fifth ran on a newer set of 120 postings (48 strong), with the incumbent run three times alongside it.
| Challenger | Its case on paper | Rank correlation, challenger vs incumbent | Strong roles | Verdict |
|---|---|---|---|---|
| 120B benchmark leader | Wins general coding and reasoning benchmarks; four times the incumbent's active parameters | 0.646 vs 0.836 to 0.853 | 11 of 12, with 10 false alarms among 32 weaker roles | Lost on ranking |
| 30B model built for speed | Looked lean in a quick demo: 231 tokens on a short test prompt | No valid result | None scored | Wrote more than 10,000 tokens per job description and ran out of room before finishing |
| Mid-size model, reasoning off | Edged the incumbent on correlation at matched settings | 0.798 vs 0.779 | 7 of 12 vs 8 to 9 | Lost on strong roles found |
| Same mid-size model, reasoning high | More thinking time per answer | 0.778 vs 0.779 | 8 of 12 vs 8 to 9 | Lost; extra reasoning bought nothing |
| Newest release, tested September 27 | Better at some hard multi-step agent tasks | 0.65 to 0.68 vs 0.71 to 0.73 | Tied on strong roles missed | Lost on ranking, and about 20 times slower |
ChallengerIncumbent range, same setup
The mid-size model’s edge is inside run-to-run noise. It also found fewer strong roles (7 of 12 vs 8 to 9), so it lost under the rule.
Rows 1 to 3: 44 job descriptions, 12 strong. Row 4: 120 postings, 48 strong. Michael Inghilterra's lab, September 2026.
Three things the scoreboard taught me
Benchmarks measure someone else's job. The 120B model wins general benchmarks and has four times the incumbent's active parameters. On my task it ranked roles far worse. A benchmark measures a task somebody else chose. Your task is the only one that counts.
Demos measure the easy case. The speed model looked lean on a short test prompt. On real job descriptions it wrote more than 10,000 tokens per answer and never finished one. A demo shows you a model on a prompt the vendor picked. Test it on your own work.
The headline number isn't the decision. The mid-size model edged the incumbent on correlation at matched settings, 0.798 to 0.779. That margin sits inside run-to-run noise, but on a vendor slide it reads as a win. It also found fewer strong roles, 7 of 12 against 8 or 9. The rule kept the incumbent, and it was right to: correlation measures the order of the whole list, and the filter's job is to not lose the best roles.
The finding that mattered more than any model
Same model, same prompt, same test set: rank correlation of 0.891 at one server setting and 0.777 at another. Two runs at the second setting agreed within 0.004, so the gap is real, not noise. I changed one configuration setting (the context window). Other things moved with it, so I haven't isolated the exact cause.
That's a swing of 0.114. It's bigger than the gap between the incumbent and every challenger except the 120B model, and more than five times the gap to the closest one. One setting moved the score more than swapping in most of the models did.
Dashed segment: range
Michael Inghilterra's lab, September 2026.
For anyone buying AI, the takeaway is blunt: the vendor's configuration isn't yours. Test in your own setup, on your own data, before you believe a number.
The incumbent isn't perfect either
Across three identical runs on the 120-posting set, the incumbent missed one strong posting in two of them, and a different posting each time. Same model, same inputs, different answers.
So I stopped treating the filter as deterministic. A random share of everything it discards now goes to the full evaluation anyway. It's a QA sample, the same habit as auditing closed-lost reasons instead of trusting the dropdown.
How I'd run this for a revenue team's AI purchase
- Write the pass/fail rule before the first demo. Two or three numbers, and at least one that matches the decision the tool makes.
- Test on your own data, including the cases that matter most, and ask for results on data the model never saw.
- Test in your configuration, not the vendor's.
- Run it more than once. If the same input gives different answers, plan an audit sample from day one.
- Keep the incumbent unless the challenger wins on the rule. "Newer" isn't a criterion.
Limits
- Small test sets. 44 roles (12 strong) for four attempts, and 120 (48 strong) for the fifth. On the 44-role set, the incumbent's own correlation moved by about 0.04 between identical runs, so treat small gaps as noise.
- The benchmark is Claude's judgment, not hiring outcomes. I measured agreement with Claude, not who gets hired.
- One task, one machine. A model that loses here can win elsewhere. The newest release really was better at some hard agent tasks. It just wasn't better at this one.
When your team last bought a tool, who wrote the pass/fail rule, and when?
Related: The fine-tune that looked fine, what happened when I tried to train the incumbent to judge like Claude. This lab runs inside trajecktory, my open-source job-search command center: trajecktory on GitHub
Keep reading