The fine-tune that looked fine
My fine-tuned AI model matched the original on the headline metric and cut false alarms by more than half. It also found 6 of the 14 roles that mattered. The original found 10.
The short version
- I fine-tuned a local AI model to judge job descriptions the way Claude does, then tested it on 116 roles it had never seen.
- On the averages, it looked like an upgrade. On the roles that mattered, it found 6 of 14. The original found 10.
- It had learned to play it safe. It never scored any role above 4.3 out of 5, so the best roles sat just under the bar.
- I wrote the pass/fail rule before I trained anything. Three versions later, none passed. The original model stays in production.
The question
I run a small AI lab at home on an NVIDIA DGX Spark. One of its jobs is triage for my job search: a local model gives incoming job descriptions a first read, and only the promising ones go on to Claude for a full evaluation.
That local model works, but it carries a hidden cost. To score like Claude, it needs my full scoring rubric in every prompt: about 22,000 tokens of instructions before it reads a single word of the job description.
Fine-tuning promised a shortcut. Train the model on Claude's past judgments, and the judgment moves into the model itself. The prompt drops from about 22,000 tokens to about 550.
So the question was simple, and it's the same one every revenue team asks about a cheaper tool: can it make the same calls as the expensive one?
The rule I wrote first
Before I trained anything, I wrote down what "better" had to mean. It's the same rule I'd applied to every model that tried to replace the one in production:
- It has to rank job descriptions at least as well as the current model (a higher rank correlation with Claude's scores).
- It has to find at least as many strong roles. No fewer.
The second condition matters most, because it matches the decision the filter makes. If the filter misses a strong role, I never see a job I'd want. If it raises a false alarm, the cost is one extra evaluation by Claude. Those two mistakes don't cost the same, so I don't count them the same.
What I built
I had 1,160 of Claude's past evaluations to learn from. Each one pairs a job description with Claude's scores and its written reasoning. A strong role is anything Claude scored 4.0 or higher out of 5.
I held back a random 116 of them, 14 of them strong, and never used them for training or tuning. Every result below comes from those 116. The model trained on the other 1,044.
I trained three versions:
- Version 1 repeated the strong roles three times in training, so the model couldn't learn to score everything low.
- Version 2 pushed harder: strong roles six times, near-strong roles twice.
- Version 3 kept version 2's recipe but trained on the scores alone. I left out Claude's written reasoning.
The first run took about six hours on one desktop box. I ran every comparison on the same serving setup, so the only thing that changed between runs was the model.
What happened
| Rank correlation | Strong roles found | False alarms | Big misses | |
|---|---|---|---|---|
| Original | 0.705 | 10 of 14 | 22 of 101 | 4 |
| Version 1 | 0.707 | 6 of 14 | 10 of 102 | 1 |
| Version 2 | 0.634 | 9 of 14 | 16 of 102 | 0 |
| Version 3 | 0.472 | 5 of 14 | 13 of 102 | 0 |
Big miss: a strong role scored 3.0 or lower. The original shows 101 non-strong roles instead of 102 because one of its answers came back malformed.
Version 1 is the one that fooled me, briefly. Its rank correlation tied the original. On a separate test set, the original's own correlation moved by 0.04 between two identical runs, so a gap of 0.002 is a tie.
It also cut false alarms by more than half. On a slide, it was an upgrade.
Then the column that decides anything: it found 6 strong roles out of 14. The original found 10.
Version 2 traded back the other way. It found 9 of 14, with more false alarms than version 1 and a weaker ranking. Both versions gave up strong roles to cut false alarms, just in different amounts. Neither found as many strong roles as the original.
116 held-out job descriptions, 14 strong. Michael Inghilterra's lab, September 2026.
Why it missed: the squeeze at the top
The fine-tuned model never gave any of the 116 roles a score above 4.3. Not once.
That's the whole failure. The bar for a strong role is 4.0. When a model's highest score is 4.3, every strong role has to land in a narrow band to clear it. Seven of the eight strong roles version 1 missed scored between 3.4 and 3.9. Close, consistently, and under the line.
The original misses differently. Its four misses were big: it scored each of those strong roles a 2 out of 5. So the honest comparison isn't a good model against a bad one. The original is erratic. The fine-tune is timid. Timid looks better on an average, because it never makes a wild call. It's worse at the one job the filter exists to do.
At or above 4.0Below 4.0
Scores for the 14 strong roles among 116 held-out job descriptions. Michael Inghilterra's lab, September 2026.
What I tried to rescue it
The obvious fix for a model that scores too low is to adjust its scores afterward. I tried four ways: shift every score up, stretch the scores to match Claude's spread, fit a straight line to Claude's scores, and lower the cutoff.
None of them held up on a second test set. The first three collapsed: the adjusted models found at most 1 of the 12 strong roles on that set. Lowering the cutoff kept the strong roles and roughly tripled the false alarms.
One option looked perfect. At a cutoff of 3.5 instead of 4.0, version 1 kept all 12 strong roles on the second set. But I found that cutoff by studying that same set, and a test you've already studied can't certify anything. I didn't count it.
The surprise: the reasoning was doing the work
Version 3 was supposed to be the efficient one. Every evaluation Claude writes includes a summary and evidence for each score, and I wondered whether all that prose was wasting the model's capacity. So I trained a version on the scores alone. Its answers ran about a quarter of the length.
It was the worst of the four. Rank correlation fell from about 0.7 to 0.472.
The likeliest reading: writing out the evidence is how the model reasons its way to a score. Take away the writing and you take away the thinking. It's the same reason I'd rather read a rep's notes on a deal than see only their close probability.
The verdict
No version beat the original. I closed the experiment on September 26, and the original model stays in production, unchanged.
I'll reopen it once I have a fresh set of about 35 strong roles the model has never seen, and roughly 600 more evaluations to train on. The next attempt keeps the reasoning and tells the model to weight the scores more heavily, instead of deleting everything around them.
What this means if you run a revenue team
I'm not an ML engineer. My career is revenue operations, sales development and analytics, and this failure is the most familiar one I know.
Averages hide the deals that matter. A forecast can land in total and miss every big deal inside it. A win rate can hold steady while you lose every enterprise deal. My model was right on average and wrong on the roles I'd actually take. Always ask what the number looks like at the top, not just across the whole list.
Pick the metric that matches the decision. The filter's job is to never lose a strong role. Correlation doesn't measure that. Strong-role recall does. If I had judged on the headline metric, I would have shipped a worse filter and felt good about it.
Write the rule before the test. Once you've seen the results, every metric starts to look like the right one. Writing the rule first is the only way to stop yourself from picking the number that flatters the answer you wanted.
Three questions to ask any AI vendor before a pilot:
- What's the hit rate on the cases that matter, not the average?
- Did you measure it on data the model never saw?
- Who wrote the pass/fail rule, and was it before or after the test?
Limits
- The sample is small. With 14 strong roles, one role moves the hit rate by about 7 points. Treat the direction as solid and the exact numbers as rough.
- The benchmark is Claude's judgment, not hiring outcomes. I measured whether the local model agrees with Claude, not whether either one predicts who gets hired.
- One task, one box. Training on a single desktop machine capped each training example at about 3,000 tokens. More memory would allow longer examples, and that might change the result.
Glossary
- Fine-tuning: training an existing model further on your own examples so it picks up your specific judgment. I used LoRA, a lightweight method that trains a small add-on instead of the whole model.
- Rank correlation: how closely two scorers agree on the order of a list, from 0 (no agreement) to 1 (the same order).
- Strong-role recall: of the roles Claude rated strong, how many the model also rated 4.0 or higher.
- False alarm: a role the model rated strong that Claude didn't.
This experiment runs inside trajecktory, the open-source job-search command center I built. It's on GitHub: trajecktory on GitHub
Keep reading