How I calibrated an LLM judge to grade like me, 25× cheaper

Evals on 188 answers from ChatGPT, Claude, Chatbase and my own pipeline over product datasheets, and a hybrid judge with Jev that grades like I do.

Businesses that sell technical products answer the same kind of question every day. “What’s the accuracy on this range?” “Can it measure through a coating, and how thick?” “Does it come with a calibration certificate?” The answers sit in product datasheets and manuals.

I took 27 of those PDFs, three of them scans with no text layer, and wrote 47 questions the way customers actually ask them. Then I had four AI systems answer every question and graded all 188 answers.

Running the systems took an afternoon. Building a judge I could trust took the rest of the work, and that is what this post is about.

The result

Correct answers out of 47: my pipeline 46, Claude app 45, ChatGPT app 35, Chatbase 33

System Correct Wrong Made up
My pipeline 46 1 0
Claude app 45 1 0
ChatGPT app 35 10 0
Chatbase 33 11 0

My pipeline is Claude Sonnet over the API with page citations. The Claude app ran Opus with the documents in a project, ChatGPT ran with thinking off, and Chatbase was on the free plan with its default model.

Every system got the same PDFs and the same instructions. One run each, so I read a one-question gap as a tie.

No system invented a spec value. I checked every claim twice: a claim checker read each one against the page it came from, and I went through the ones it was unsure about by hand. The weaker systems failed more quietly. They said an answer wasn’t in the documents when it was.

Where the misses came from

Correct answers by question type for each system

Scanned pages. Systems that only read the text layer saw blank pages. Every question answered only by a scan came back “not specified”.

Answers that need a second look. A product advertises one range on the front page. A different mode of the same product stops at a tenth of it, and the customer’s question is about that mode. The same pattern showed up in accuracy tables split by range and in specs that change with the material.

Unit conversions. The customer asks in one unit. The datasheet lists the other unit in one row and a separate, lower limit in the customer’s unit in another. Converting the first number gives a confident wrong answer.

The website and the datasheet disagreeing. The product page said one value and the datasheet said half of it. My rule: the datasheet wins, and the reply says the page may be wrong. Two systems missed it.

Rules that live in no PDF

Before running anything I reviewed every question. I dropped three that no customer would ask, corrected two gold answers, and turned the “not in the documents” cases into rules:

After grading I added one more: answer what was asked. Extra conditions confuse buyers and invite more questions.

That short list did more for answer quality than any prompt trick.

Why the judge needs a judge

Grading 188 answers by hand takes hours, so the usual move is an LLM judge: give a model the question, the gold answer and the answer, and ask for correct, partial or wrong. You only know the judge is any good if it agrees with the person whose standard matters. Here, that’s me.

So I graded 20 answers blind: no system names, five per system, with some likely mistakes mixed in so the judge would be tested on errors too. Then I compared.

First try: 13 of 20. Most of the gap was in my own answer key. Three gold answers were wrong. Grading real answers showed me that I answer from the datasheet as printed, even when a conversion on it looks off, and flag the document separately. My key hadn’t been written that way. Once I fixed it, agreement jumped.

Plain agreement flatters a judge when most answers are correct. A judge that marks everything “correct” would agree 80% of the time here and understand nothing. Cohen’s kappa corrects for that.

Cohen’s kappa: 67 of 100 agreement expected by luck, 28 from skill, 5 missed; kappa = 28 / 33 = 0.85

Kappa asks how far above luck the judge got, as a share of how far above luck it could have got. 0 means no better than random; 1 means it matched me every time.

Two judges

Agreement with my grades: Sonnet 35 of 40, hybrid 36 of 40. Cost for 188 answers: $0.80 vs $0.03

Judge Agrees Cost, 188 Speed
Sonnet alone 35 / 40 ~$0.80 seconds
Hybrid 36 / 40 ~$0.03 under 1 s

The hybrid gives each part the job it does best.

The hybrid judge: code compares numbers, Jev answers typed questions, code applies the policy; Sonnet with the PDF only for unsure claims

Code pulls the numbers out of the gold answer and the answer being graded, and records which match. It never confuses ±(1.2% + 5) with ±(0.8% + 5).

Jev handles meaning. It’s a System One model from TypeSafe: it reads natural language and returns typed answers with probabilities, with no prose to parse. One request per answer asks four questions at once:

{
  "verdict": {"type": "choice",
    "instructions": "Grade `answer` to `customer_question` against `gold` and `notes`. Use `code_checks` for which gold numbers the answer contains.",
    "criteria": {"correct": "...", "partial": "...", "wrong": "..."}},
  "brush_off": {"type": "noul",
    "instructions": "Does `answer` brush the customer off, such as a bare 'not in our documents' with no next step?"},
  "unasked_extra": {"type": "noul",
    "instructions": "Does `answer` add specs, conditions or other models that `customer_question` didn't ask about?"},
  "rule_followed": {"type": "noul",
    "instructions": "Does `answer` do what `rule` says?"}
}

A Choice picks one of the grades and returns a probability for each. A Noul is the probability that a statement holds. The judgment comes back as numbers, so the policy stays in code where I can see it and change it. One real answer from the run:

verdict        correct 0.98 · partial 0.02 · wrong 0.00
unasked_extra  0.91  → at or above 0.9: "too much information" flag

About 2,000 tokens per answer, under a second, and all 188 answers for about three cents.

Sonnet with the original PDF only sees the few claims Jev isn’t sure about.

Keeping myself honest

I wrote the pass line down before I graded a second set of 20: at least 16 of 20, kappa at least 0.6. That second set got used once, as a test, and never for tuning.

It caught a mistake. On the first 20, I had tuned a penalty: if Jev was 90% sure an answer added unasked details, the grade dropped to partial. It looked perfect there. On the fresh set it caused two of the three disagreements. The judge passed the line (17 of 20, kappa 0.63), but barely.

Reading those two answers showed why. I grade substance and length separately. An answer can be right and too long. The detector itself was right: it flagged exactly the four answers across both sets that I had found too long, and no others. The penalty was the wrong policy. So it became a flag beside the grade. Without it the judge agrees with me 19 of 20, kappa 0.85. That change came after I’d seen the second set, so the next fresh set is where it has to hold.

The loop I’d run every time

The eval loop: golden set, all systems answer, I grade 20 blind, the judge grades the same 20 and I fix each disagreement, a fresh 20 against a line set in advance, then grade everything

I did it in a worse order: ran the systems, graded everything with an untested judge, reported numbers, then calibrated. Next time:

  1. Build the golden set: questions as asked, reviewed gold answers.
  2. All systems answer, same documents, same instructions.
  3. I grade 20 answers blind.
  4. The judge grades the same 20. I read every disagreement and fix the cause, which is usually the answer key.
  5. Write the pass line down, then test on a fresh 20. Use it once.
  6. Only then grade everything and report.

A golden set isn’t finished until a person has checked real answers against it.

What it means

On product datasheets, a frontier model with the PDFs attached is already very good: my pipeline and the Claude app tied. Accuracy is the starting point. The hard part is getting that accuracy into the place where buyers ask, across a catalogue too big for one prompt, while datasheets change and the website drifts away from them. Then proving it on the business’s own questions, with a judge calibrated to the person who answers them. That is what I’m building now.

The whole run cost about $8 in API calls. Most of that went on mistakes I won’t repeat. Now I estimate every paid run from a one-question trial, counting every path that can call a paid model.