XLNCXLNC Watch it measure

Blog

Evaluating the evaluator: measure an LLM judge before you trust its score

An LLM judge's score is only as good as the judge. Here is what it means for a judge to be calibrated, why judge drift differs from model drift, and where AIM's judge credentialing stands today: credentials at levels 6 to 10, with re-credentialing under a stricter procedure under way.

An LLM graded your model's answer a 4 out of 5. You shipped on that number, or you did not.

Here is the question almost nobody in the pipeline can answer: how good is the grader?

Short answer. An LLM judge is an instrument, and an instrument has to be measured before its readings count. A calibrated judge has a known severity at each level of the work, a fit to the scale checked against a rule declared in advance, and a standard error on every score. AIM credentials judges level by level through a formal admission gate. That procedure is designed and partly piloted. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way.

The grader is also a measurement

When a model scores another model's output, two measurements are happening at once. One is of the answer. The other, usually unexamined, is of the judge: its tastes, its blind spots, the levels of difficulty where it reads well and the levels where it guesses.

Claim. LLM judges carry systematic biases that shift their scores independent of answer quality.

Evidence. Zheng et al. (2023), "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," NeurIPS 2023 Datasets and Benchmarks track, document position bias, verbosity bias, and self-enhancement bias in LLM judges.

None of this makes LLM judges useless. It makes them instruments. A thermometer that reads two degrees warm is still useful once you know it reads two degrees warm. The trouble starts when nobody checks.

What "calibrated" means here

Four ideas carry the whole approach. Each one is plain.

Severity and leniency

Some judges grade hard, some grade soft, and the lean often changes with the difficulty of the work. Severity is where the judge sits on the same scale as the work it scores, disclosed with its standard error. A severity value is a disclosure about bias. It is not a grade for quality.

Fit

Severity only means something where the judge's scores behave consistently with the scale. Fit is checked against a band declared before the data is analyzed, so nobody can move the goalposts after seeing the result.

Standard error

Every score a calibrated judge produces carries its own uncertainty. A 4 with a standard error of 0.1 and a 4 with a standard error of 1.0 support very different decisions.

Measured range

A judge is calibrated for the levels where it was measured and passed. Outside that range it is excluded for the item, not trusted by default.

Evidence, early and not yet certified. We published a pilot calibration of one of the judges we use in Use a judge only where it was measured. It was measured at eight levels of task complexity against a fit band of 0.50 to 1.50, declared before the fit was computed. Two levels cleared that fit band. Six did not, and we published those too. Calibration vintage: 2026-09-20; a later re-analysis under a revised gate admitted more levels, most only at whole-stage grain.

The method behind this treats each judge as a rater on one shared scale, so its lean can be compared directly with human graders and other judges. It implements at measurement grade the latent-trait/GLMM approach that NIST AI 800-3 highlights as a promising foundation for AI evaluation statistics.

Why averaging three judges is not the same thing

A common shortcut is to run three judges and average them. It is cheaper than calibration, and it does reduce random disagreement.

What it cannot do is remove a lean the judges share. If all three grade leniently on hard items, the average grades leniently on hard items too, and it gives no statement of its own error. You get a smoother number with the same bias and no error bar.

Calibration measures each judge's lean and precision at each level, so a lean can be corrected for, and a judge that does not fit a level can be left out of it. If your eval platform already runs judges, keep them. Calibration is a measurement of those judges, not a replacement for the tool that runs them.

Judge drift is not model drift

Most drift monitoring watches the system being evaluated. That matters, and it is well covered elsewhere.

Judge drift is the quieter problem. Many judges are vendor-served models, and a provider release can change a judge's behavior under the same name. A judge that fit the scale at level 8 last month may not fit it this month, and nothing in the score will tell you. The scores keep arriving. They just mean something different.

The remedy is to treat a judge's calibration as perishable. In our pilot, a judge's entries go stale after 90 days or a version bump, whichever comes first, and stale entries are marked rather than deleted. A stale judge gets re-measured before its scores count again.

Credentialing: a formal admission gate

Credentialing turns calibration into a rule. Under AIM's procedure, a judge is admitted level by level, not with one overall grade. At each level, three conditions must hold: the judge's fit stays inside the declared band, its precision meets the target, and a lower-bound check on its performance at that level passes. A judge that clears some levels and not others is credentialed for the levels it cleared.

That per-level shape matters. A single overall credential would hide exactly the levels where a judge fails.

Checking judges is not a new idea, and some tools already help. LangSmith's Align Evals calibrates a judge against human feedback, and Patronus reports judge-to-human agreement. Both are useful checks. What we have not seen described in public materials is a credential issued or withheld per capability level, against bands declared before the data came in. That gate is the part AIM adds.

What exists today, and what does not

ComponentStatus (early evidence; not yet certified)
Rater-scale calibration of a judgeTested early; one judge's results published, vintage 2026-09-20.
Admission gate (fit, precision, lower bound, per level)Designed; passes in design review. Not yet certified.
Judge lifecycle rules (staleness, re-measurement)Draft, awaiting certification.
Credentials issued under the full procedureNone yet. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way.
Coverage at the highest complexity levelsNot started. We make no claim at those levels.

Stated plainly: we can show how a judge is measured, and we have published the results of doing it once, failures included. We have not yet issued a credential under the full procedure, and we will not describe any judge as credentialed under it until one is.

Frequently asked questions

What does it mean for an LLM judge to be calibrated?

A calibrated judge has been measured on the same scale as the work it scores. Its severity or leniency at each level is known, its fit to the scale has been checked against a rule declared in advance, and every score it produces carries a standard error. A judge is calibrated for the levels where it was measured, not in general.

How is an LLM judge credentialed in AIM?

AIM's credentialing procedure admits a judge level by level. At each level the judge's fit must hold under a pre-declared rule, its precision must meet a target, and a lower-bound check on its performance must pass. The procedure is designed and parts of it have been piloted. It has not yet been certified, and no credentials have been issued under it.

How is judge drift different from model drift?

Model drift is a change in the system being evaluated. Judge drift is a change in the grader. A vendor-served judge can change behavior under the same name after a provider release, so a judge that fit the scale last month may not fit it now. Judge drift calls for re-measuring the judge, not the model.

Is averaging several LLM judges the same as calibrating them?

No. Averaging reduces random disagreement between judges, but if the judges share a lean at a given level, the average carries that lean, and it does not state its own error. Calibration measures each judge's lean and precision at each level so they can be corrected for or excluded.

Does AIM have credentialed judges today?

Not under the full procedure yet. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way. XLNC has also published a pilot calibration of one judge across eight levels of task complexity.

Where this connects

A calibrated judge is also what makes fresh test items trustworthy: a new item located by an uncalibrated judge inherits that judge's lean. That is the subject of Fresh instruments on demand. For how capability and severity interact, see The two numbers that decide whether you can trust a model.

If you run LLM judges today and want to know where they can be trusted, join the waitlist. Early access is by conversation.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.