Calibrating a Model Judge — Diagram
How to establish what a judge's verdicts are worth before relying on them — the human-labelled set, the agreement measurement, and the drift check that catches the judge changing underneath you.
SVG. No sign-up, no email.
Using a model to judge another model's output is the only practical way to evaluate at scale. It is also an instrument with an unknown error rate until you measure it — and almost nobody measures it.
The sequence below is the calibration. It costs one afternoon and it converts every number the judge produces afterwards from an assertion into a measurement with a stated confidence.
The number that matters#
Agreement with human labels. If the judge agrees 80% of the time, every figure it produces afterwards carries roughly that error, and a reported "94% groundedness" should be understood as 94% ± the judge's own uncertainty.
That does not make the judge useless. It makes it an instrument with a known error bar, which is the difference between a measurement and a claim.
Human disagreement is a finding, not a nuisance#
The step where two annotators disagree and resolve it is shaded as a warning for a reason: the rate of human disagreement is a ceiling on the judge.
If two experienced people disagree on 15% of cases, no judge can be expected to exceed 85% agreement, and chasing that is wasted effort. It also tells you something more useful — that the criterion itself is ambiguous, and the fix is a clearer rubric rather than a better judge.
Systematic versus scattered disagreement#
Scattered disagreement is noise and it is tolerable.
Systematic disagreement is a bias you can correct. If the judge consistently marks long answers as better, or consistently accepts a claim that is plausible but unsupported, that is a fixable prompt problem — and it is invisible unless you look at where the disagreements fall rather than just counting them.
Why the last lane has no arrows#
Those four are standing practices, not a sequence. The critical pairing is pin the judge model and keep the reference set unchanged: if either moves, your historical comparisons quietly stop meaning anything, and the trend line you have been tracking for six months becomes fiction.
The narrower the question, the better the judge#
Calibration is easier and agreement higher when the judge is asked something specific — "does this answer contain a claim not supported by the supplied passages?" — rather than "is this a good answer?"
If agreement comes out low, the first thing to try is a narrower question, not a larger model.
See model evaluation for the full method, groundedness for what a judge is typically asked to check, and AI Testing for why we publish this rather than sell a tool.