← All writing
Lab notes

The fine-tune that looked fine

My fine-tuned AI model matched the original on the headline metric and cut false alarms by more than half. It also found 6 of the 14 roles that mattered. The original found 10.

Michael InghilterraLab notesSeptember 20267 min read
airevopsmeasurementlab-notes
0.707 vs 0.705Rank correlation with Claude. A tie.
10 vs 22False alarms. Better.
6 vs 10 of 14Strong roles found. Worse.

The short version

  • I fine-tuned a local AI model to judge job descriptions the way Claude does, then tested it on 116 roles it had never seen.
  • On the averages, it looked like an upgrade. On the roles that mattered, it found 6 of 14. The original found 10.
  • It had learned to play it safe. It never scored any role above 4.3 out of 5, so the best roles sat just under the bar.
  • I wrote the pass/fail rule before I trained anything. Three versions later, none passed. The original model stays in production.

The question

I run a small AI lab at home on an NVIDIA DGX Spark. One of its jobs is triage for my job search: a local model gives incoming job descriptions a first read, and only the promising ones go on to Claude for a full evaluation.

That local model works, but it carries a hidden cost. To score like Claude, it needs my full scoring rubric in every prompt: about 22,000 tokens of instructions before it reads a single word of the job description.

Fine-tuning promised a shortcut. Train the model on Claude's past judgments, and the judgment moves into the model itself. The prompt drops from about 22,000 tokens to about 550.

So the question was simple, and it's the same one every revenue team asks about a cheaper tool: can it make the same calls as the expensive one?

The rule I wrote first

Before I trained anything, I wrote down what "better" had to mean. It's the same rule I'd applied to every model that tried to replace the one in production:

  1. It has to rank job descriptions at least as well as the current model (a higher rank correlation with Claude's scores).
  2. It has to find at least as many strong roles. No fewer.

The second condition matters most, because it matches the decision the filter makes. If the filter misses a strong role, I never see a job I'd want. If it raises a false alarm, the cost is one extra evaluation by Claude. Those two mistakes don't cost the same, so I don't count them the same.

What I built

I had 1,160 of Claude's past evaluations to learn from. Each one pairs a job description with Claude's scores and its written reasoning. A strong role is anything Claude scored 4.0 or higher out of 5.

I held back a random 116 of them, 14 of them strong, and never used them for training or tuning. Every result below comes from those 116. The model trained on the other 1,044.

I trained three versions:

  • Version 1 repeated the strong roles three times in training, so the model couldn't learn to score everything low.
  • Version 2 pushed harder: strong roles six times, near-strong roles twice.
  • Version 3 kept version 2's recipe but trained on the scores alone. I left out Claude's written reasoning.

The first run took about six hours on one desktop box. I ran every comparison on the same serving setup, so the only thing that changed between runs was the model.

What happened

Rank correlationStrong roles foundFalse alarmsBig misses
Original0.70510 of 1422 of 1014
Version 10.7076 of 1410 of 1021
Version 20.6349 of 1416 of 1020
Version 30.4725 of 1413 of 1020

Big miss: a strong role scored 3.0 or lower. The original shows 101 non-strong roles instead of 102 because one of its answers came back malformed.

Version 1 is the one that fooled me, briefly. Its rank correlation tied the original. On a separate test set, the original's own correlation moved by 0.04 between two identical runs, so a gap of 0.002 is a tie.

It also cut false alarms by more than half. On a slide, it was an upgrade.

Then the column that decides anything: it found 6 strong roles out of 14. The original found 10.

Version 2 traded back the other way. It found 9 of 14, with more false alarms than version 1 and a weaker ranking. Both versions gave up strong roles to cut false alarms, just in different amounts. Neither found as many strong roles as the original.

No version found as many strong roles as the original
Scatter plot of four model versions. The original found 10 of 14 strong roles with 22 false alarms. Version 1 found 6 with 10 false alarms, version 2 found 9 with 16, and version 3 found 5 with 13. No version found as many strong roles as the original.024681012140510152025Strong roles found, of 14 (higher is better)False alarms (lower is better)BetterVersion 1Version 2Version 3Original

116 held-out job descriptions, 14 strong. Michael Inghilterra's lab, September 2026.

Why it missed: the squeeze at the top

The fine-tuned model never gave any of the 116 roles a score above 4.3. Not once.

That's the whole failure. The bar for a strong role is 4.0. When a model's highest score is 4.3, every strong role has to land in a narrow band to clear it. Seven of the eight strong roles version 1 missed scored between 3.4 and 3.9. Close, consistently, and under the line.

The original misses differently. Its four misses were big: it scored each of those strong roles a 2 out of 5. So the honest comparison isn't a good model against a bad one. The original is erratic. The fine-tune is timid. Timid looks better on an average, because it never makes a wild call. It's worse at the one job the filter exists to do.

The fine-tune's scores bunch just under the bar
Dot plot of the scores two models gave the same 14 strong roles. The original scored 10 of them at 4.0 or above, spread up to 5.0, but scored four of them at 2.0. The fine-tuned model never scored above 4.3, and seven of its eight misses landed between 3.4 and 3.9, just under the 4.0 bar.Bar for a strong role (4.0)2.03.04.05.0Score given (out of 5)Original modelFour big missesFine-tuned (version 1)Never scoredabove 4.3

At or above 4.0Below 4.0

Scores for the 14 strong roles among 116 held-out job descriptions. Michael Inghilterra's lab, September 2026.

What I tried to rescue it

The obvious fix for a model that scores too low is to adjust its scores afterward. I tried four ways: shift every score up, stretch the scores to match Claude's spread, fit a straight line to Claude's scores, and lower the cutoff.

None of them held up on a second test set. The first three collapsed: the adjusted models found at most 1 of the 12 strong roles on that set. Lowering the cutoff kept the strong roles and roughly tripled the false alarms.

One option looked perfect. At a cutoff of 3.5 instead of 4.0, version 1 kept all 12 strong roles on the second set. But I found that cutoff by studying that same set, and a test you've already studied can't certify anything. I didn't count it.

The surprise: the reasoning was doing the work

Version 3 was supposed to be the efficient one. Every evaluation Claude writes includes a summary and evidence for each score, and I wondered whether all that prose was wasting the model's capacity. So I trained a version on the scores alone. Its answers ran about a quarter of the length.

It was the worst of the four. Rank correlation fell from about 0.7 to 0.472.

The likeliest reading: writing out the evidence is how the model reasons its way to a score. Take away the writing and you take away the thinking. It's the same reason I'd rather read a rep's notes on a deal than see only their close probability.

The verdict

No version beat the original. I closed the experiment on September 26, and the original model stays in production, unchanged.

I'll reopen it once I have a fresh set of about 35 strong roles the model has never seen, and roughly 600 more evaluations to train on. The next attempt keeps the reasoning and tells the model to weight the scores more heavily, instead of deleting everything around them.

What this means if you run a revenue team

I'm not an ML engineer. My career is revenue operations, sales development and analytics, and this failure is the most familiar one I know.

Averages hide the deals that matter. A forecast can land in total and miss every big deal inside it. A win rate can hold steady while you lose every enterprise deal. My model was right on average and wrong on the roles I'd actually take. Always ask what the number looks like at the top, not just across the whole list.

Pick the metric that matches the decision. The filter's job is to never lose a strong role. Correlation doesn't measure that. Strong-role recall does. If I had judged on the headline metric, I would have shipped a worse filter and felt good about it.

Write the rule before the test. Once you've seen the results, every metric starts to look like the right one. Writing the rule first is the only way to stop yourself from picking the number that flatters the answer you wanted.

Three questions to ask any AI vendor before a pilot:

  1. What's the hit rate on the cases that matter, not the average?
  2. Did you measure it on data the model never saw?
  3. Who wrote the pass/fail rule, and was it before or after the test?

Limits

  • The sample is small. With 14 strong roles, one role moves the hit rate by about 7 points. Treat the direction as solid and the exact numbers as rough.
  • The benchmark is Claude's judgment, not hiring outcomes. I measured whether the local model agrees with Claude, not whether either one predicts who gets hired.
  • One task, one box. Training on a single desktop machine capped each training example at about 3,000 tokens. More memory would allow longer examples, and that might change the result.

Glossary

  • Fine-tuning: training an existing model further on your own examples so it picks up your specific judgment. I used LoRA, a lightweight method that trains a small add-on instead of the whole model.
  • Rank correlation: how closely two scorers agree on the order of a list, from 0 (no agreement) to 1 (the same order).
  • Strong-role recall: of the roles Claude rated strong, how many the model also rated 4.0 or higher.
  • False alarm: a role the model rated strong that Claude didn't.

This experiment runs inside trajecktory, the open-source job-search command center I built. It's on GitHub: trajecktory on GitHub

Keep reading

Michael Inghilterra
Michael Inghilterra
RevOps & Analytics · Sales Development · building trajecktory
← All writing