XLNCXLNC Watch it measure

Blog

Our judges' credentials held on a second measurement

We paid to measure three judges twice. Here is what moved, and what it does not prove.

The question

Your eval score moved. Was it the model, or the judge? If you never measured the judge twice, you cannot say.

So we measured our own. We re-measured three LLM judges on the same calibration items in a second paid run: Qwen3 235B Instruct (qwen3-235b-a22b-instruct-2507), DeepSeek V4.1 Flash (deepseek-v4-1-flash) and TypeSafe (Jev).

What we found

Test-retest reliability of judge ratings (ICC(2,1), items centered within MHC level): DeepSeek V4.1 Flash 0.993 (95% CI 0.990 to 0.995, 113 items), Qwen3 235B 0.967 (95% CI 0.954 to 0.976, 137 items), TypeSafe 0.999 (95% CI 0.998 to 0.999, 113 items).

Average judge severity drift between the two runs was under 0.01 logit for every judge (DeepSeek +0.005, Qwen3 +0.007, TypeSafe -0.001; standard error about 0.03), statistically equivalent to zero within a pre-declared margin of plus or minus 0.10 logit.

All 19 retested judge-by-level cells (MHC levels 6 to 12) matched their certified calibration within two standard errors, and none showed misfit above the 1.5 mean-square limit.

The tests and the margin were written down before the second run started. The chart for each judge is on the Judges page.

What this does not show

These judges ran at temperature 0, so much of this stability reflects near-deterministic model output; the figures describe LLM judge stability on a fixed item bank at stage-level grain, not the reliability of scores for people.

The two runs were close together. Whether the same judges hold steady over months is a separate check, and it is not done yet.

What it means for you

When a judge's severity is measured and shown to hold, a change in your score is more likely to be a change in your model. When it is not measured, you are guessing. Next we will check the same judges over a longer interval and publish that result whether it holds or not.

See every judge's record.

The background, and why a changed judge can pass for a better product, is in Your judge changed. Did your scores notice?

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.