How I calibrated an LLM judge to grade like me, 25× cheaper
Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do.
Businesses that sell technical products answer the same kind of question every day. “What’s the accuracy on this range?” “Can it measure through a coating, and how thick?” “Does it come with a calibration certificate?” The answers sit in product datasheets and manuals.
I took 27 of those PDFs, three of them scans with no text layer, and wrote 47 questions the way customers actually ask them. Then I had four AI systems answer every question and graded all 188 answers.
Running the systems took an afternoon. Building a judge I could trust took the rest of the work, and that is what this post is about.
The result

| System | Correct | Wrong | Made up |
|---|---|---|---|
| My pipeline | 46 | 1 | 0 |
| Claude app | 45 | 1 | 0 |
| ChatGPT app | 35 | 10 | 0 |
| Chatbase | 33 | 11 | 0 |
My pipeline is Claude Sonnet over the API with page citations. The Claude app ran Opus with the documents in a project, ChatGPT ran with thinking off, and Chatbase was on the free plan with its default model.
Every system got the same PDFs and the same instructions. One run each, so I read a one-question gap as a tie.
No system invented a spec value. I checked every claim twice: a claim checker read each one against the page it came from, and I went through the ones it was unsure about by hand. The weaker systems failed more quietly. They said an answer wasn’t in the documents when it was.
Where the misses came from

Scanned pages. Systems that only read the text layer saw blank pages. Every question answered only by a scan came back “not specified”.
Answers that need a second look. A product advertises one range on the front page. A different mode of the same product stops at a tenth of it, and the customer’s question is about that mode. The same pattern showed up in accuracy tables split by range and in specs that change with the material.
Unit conversions. The customer asks in one unit. The datasheet lists the other unit in one row and a separate, lower limit in the customer’s unit in another. Converting the first number gives a confident wrong answer.
The website and the datasheet disagreeing. The product page said one value and the datasheet said half of it. My rule: the datasheet wins, and the reply says the page may be wrong. Two systems missed it.
Rules that live in no PDF
Before running anything I reviewed every question. I dropped three that no customer would ask, corrected two gold answers, and turned the “not in the documents” cases into rules:
- Datasheet beats product page. Give the datasheet value and say the page may have an error.
- Never a bare “not in our documents”. Say it isn’t in the published documents, and give an email to ask.
- Certifications only if the document says so. Otherwise, email for confirmation.
- Canned answers for the two most common questions.
After grading I added one more: answer what was asked. Extra conditions confuse buyers and invite more questions.
That short list did more for answer quality than any prompt trick.
Why the judge needs a judge
Grading 188 answers by hand takes hours, so the usual move is an LLM judge: give a model the question, the gold answer and the answer, and ask for correct, partial or wrong. You only know the judge is any good if it agrees with the person whose standard matters. Here, that’s me.
So I graded 20 answers blind: no system names, five per system, with some likely mistakes mixed in so the judge would be tested on errors too. Then I compared.
First try: 13 of 20. Most of the gap was in my own answer key. Three gold answers were wrong. Grading real answers showed me that I answer from the datasheet as printed, even when a conversion on it looks off, and flag the document separately. My key hadn’t been written that way. Once I fixed it, agreement jumped.
Plain agreement flatters a judge when most answers are correct. A judge that marks everything “correct” would agree 80% of the time here and understand nothing. Cohen’s kappa corrects for that.

Kappa asks how far above luck the judge got, as a share of how far above luck it could have got. 0 means no better than random; 1 means it matched me every time.
Two judges

| Judge | Agrees | Cost, 188 | Speed |
|---|---|---|---|
| Sonnet alone | 35 / 40 | ~$0.80 | seconds |
| Hybrid | 36 / 40 | ~$0.03 | under 1 s |
The hybrid gives each part the job it does best.

Code pulls the numbers out of the gold answer and the answer being graded, and records which match. It never confuses ±(1.2% + 5) with ±(0.8% + 5).
Jev handles meaning. It’s a System One model from TypeSafe: it reads natural language and returns typed answers with probabilities, with no prose to parse. One request per answer asks four questions at once:
{
"verdict": {"type": "choice",
"instructions": "Grade `answer` to `customer_question` against `gold` and `notes`. Use `code_checks` for which gold numbers the answer contains.",
"criteria": {"correct": "...", "partial": "...", "wrong": "..."}},
"brush_off": {"type": "noul",
"instructions": "Does `answer` brush the customer off, such as a bare 'not in our documents' with no next step?"},
"unasked_extra": {"type": "noul",
"instructions": "Does `answer` add specs, conditions or other models that `customer_question` didn't ask about?"},
"rule_followed": {"type": "noul",
"instructions": "Does `answer` do what `rule` says?"}
}
A Choice picks one of the grades and returns a probability for each. A Noul is the probability that a statement holds. The judgment comes back as numbers, so the policy stays in code where I can see it and change it. One real answer from the run:
verdict correct 0.98 · partial 0.02 · wrong 0.00
unasked_extra 0.91 → at or above 0.9: "too much information" flag
About 2,000 tokens per answer, under a second, and all 188 answers for about three cents.
Sonnet with the original PDF only sees the few claims Jev isn’t sure about.
Keeping myself honest
I wrote the pass line down before I graded a second set of 20: at least 16 of 20, kappa at least 0.6. That second set got used once, as a test, and never for tuning.
It caught a mistake. On the first 20, I had tuned a penalty: if Jev was 90% sure an answer added unasked details, the grade dropped to partial. It looked perfect there. On the fresh set it caused two of the three disagreements. The judge passed the line (17 of 20, kappa 0.63), but barely.
Reading those two answers showed why. I grade substance and length separately. An answer can be right and too long. The detector itself was right: it flagged exactly the four answers across both sets that I had found too long, and no others. The penalty was the wrong policy. So it became a flag beside the grade. Without it the judge agrees with me 19 of 20, kappa 0.85. That change came after I’d seen the second set, so the next fresh set is where it has to hold.
The loop I’d run every time

I did it in a worse order: ran the systems, graded everything with an untested judge, reported numbers, then calibrated. Next time:
- Build the golden set: questions as asked, reviewed gold answers.
- All systems answer, same documents, same instructions.
- I grade 20 answers blind.
- The judge grades the same 20. I read every disagreement and fix the cause, which is usually the answer key.
- Write the pass line down, then test on a fresh 20. Use it once.
- Only then grade everything and report.
A golden set isn’t finished until a person has checked real answers against it.
What it means
On product datasheets, a frontier model with the PDFs attached is already very good: my pipeline and the Claude app tied. Accuracy is the starting point. The hard part is getting that accuracy into the place where buyers ask, across a catalogue too big for one prompt, while datasheets change and the website drifts away from them. Then proving it on the business’s own questions, with a judge calibrated to the person who answers them. That is what I’m building now.
The whole run cost about $8 in API calls. Most of that went on mistakes I won’t repeat. Now I estimate every paid run from a one-question trial, counting every path that can call a paid model.