XLNCXLNC Watch it measure

Blog

The same judge, one level apart: LLM judge quality is a profile, not a score

We measured five LLM judges at eight levels of reasoning complexity. Their fit and where they qualified changed from level to level, and price did not predict either. Here are the early numbers, and three figures to ask any judge vendor for.

The answer first

An LLM judge is a different instrument at each level of reasoning complexity. In our pilot data, Jev, one of the five judges we tested, fit our measurement model as expected at level 8 of the Model of Hierarchical Complexity (MHC), with an infit of 0.97. One level down, at level 7, the same judge read 0.13.

Those two numbers describe one judge, one scale and one week of calibration. A single quality score for that judge would average them into something that is true at neither level.

We ran five judges at levels 5 through 12, 40 judge-by-level cells in all. Thirty-seven of those cells gave usable estimates; the other three, all at level 5, did not. Twenty-one cells were admitted. Seventeen of those 21 were admitted only at coarse, whole-stage grain, because the judge's ratings were too predictable to trust at finer grain. No judge was admitted at level 5, 11 or 12, and none was run as a judge at level 13. Everything in this post is early evidence, not yet certified.

How we measured: five judges, levels 5 to 12, one ruler, one measurement model (a Rasch model for part-credit answers with a judge severity term, in logits), calibrations dated 2026-09-20 (Jev) and 2026-09-23 (the other four). "Admitted" means a judge passed our admission gate at that level; it is not a verdict on the model in general. Model names identify what we tested; there is no affiliation or endorsement. XLNC has no commercial relationship with the model vendors named. An earlier post, Use a judge only where it was measured (2026-09-22), reported Jev under our previous item labels and admission rule, which admitted it at levels 8 and 10 only. This post applies the revised item labels and admission gate adopted on 2026-09-23 to the same calibration, so Jev's admitted levels, fit and severity values differ from that post.

One judge fits the model at level 8 and overfits it at every other level (by outfit)

Fit is how well a judge's ratings behave like measurement. The target is 1.0. Well below it, the judge is too predictable: it echoes a pattern, such as the rubric wording or the anchor texts, and adds little independent information of its own. Well above it, the judge is noisy. This is the measurement sense of "overfit", and it is related to, but not the same as, a model memorizing its training data.

Jev's outfit at level 8 was 1.00. At every other level we measured, its outfit sat between 0.17 and 0.35, under the usual floor of 0.5.

Five panels, one per judge, showing infit (circles) and outfit (squares) mean squares on a log scale for levels 5 to 12, with the 0.5 to 1.5 band shaded and a dashed line at 0.40. Most admitted levels fall below the band, meaning the ratings are more predictable than the model expects. Jev at level 8 is the one cell near 1.0. Early evidence, not yet certified.

The pattern is not special to Jev. Most admitted cells across all five judges sit in the overfit zone, and one model, gpt-oss-120b, fit well at levels 6 and 8, where it was not admitted, and overfit at 7 and 9, where it was. Admission weighs precision, stage separation and bias checks as well as fit, so a good fit number alone does not earn a level, and a bad one does not always lose it.

A reading of 0.01 looks like near-perfect consistency. It is closer to a gauge that shows the same number whatever you put on it. Selling that as precision overstates what the instrument knows.

The priciest judge was not admitted at level 6, where three cheaper judges were

Price is the shortcut most teams use to choose a judge. Our data does not reward it.

The priciest judge in our pilot, Kimi-k3 (Moonshot AI), cost USD 29.75 per 1,336 ratings on our measured ledger. At level 6 it was not admitted, because it did not pass our stage-separation check. That is a result of our admission gate at one level, not a verdict on the model. At the same level, three cheaper judges were admitted, two of them only at whole-stage grain: Jev at USD 0.00 (no per-call charge recorded), qwen3-235b at USD 1.16 and deepseek-v4.1-flash at USD 4.41.

A scatter of five judges: cost in US dollars per 1,336 ratings (horizontal, 0 to about 30) against the number of MHC levels admitted (vertical). Jev costs 0.00 on the ledger; Kimi-k3 costs 29.75 and is admitted at four levels, fewer than two much cheaper judges admitted at five. With five points, no relation is claimed. Early evidence, not yet certified.

Across all five, the cheapest priced judge, gpt-oss-120b at USD 0.50, was admitted at two levels, and the most expensive was admitted at four. Five points cannot establish a relation between price and coverage in either direction. They are enough to show that price does not tell you where a judge works.

Admission also has holes. gpt-oss-120b is admitted at levels 7 and 9 and excluded at 8, between them. An adaptive test that climbs levels needs a qualified judge at every level it may visit, and a hole in the middle stalls it or pushes it onto a judge nobody qualified.

A grid titled 'Price does not tell you where a judge works.' Five judges, cheapest first (Jev US$0.00, gpt-oss-120b US$0.50, qwen3-235b US$1.16, deepseek-v4.1-flash US$4.41, Kimi-k3 US$29.75, each per 1,336 ratings), by MHC levels 5 to 12. A check mark means admitted, a half-filled circle means admitted at coarse whole-stage grain only, and a cross means not admitted. No judge is admitted at levels 5, 11 or 12, and the most expensive judge is admitted at fewer levels than the cheapest. Calibrations from an earlier procedure, dated 2026-09-20 and 2026-09-23.

A judge's reasoning ceiling is not its judging credential

Here is the point that surprised us most. Measured as a performer, doing the reasoning itself, Kimi-k3 answered at the top of our earlier comprehension ruler. That ruler cannot reliably tell levels 12 and 13 apart, so we have retired it for that purpose and make no claim about where the model's own reasoning tops out. Measured as a judge, rating other people's reasoning, it was not admitted at levels 5, 6, 11 and 12; at level 5 we could not estimate it at all. At level 11 its ratings fell outside our fit range on the noisy side, with an infit of 1.62 and an outfit of 2.42. It was admitted only at levels 7 through 10, and even there its infit ran from 0.01 to 0.18.

Doing a task and grading it are different skills. A model that can produce top-level reasoning has not thereby shown that it can tell level 11 work from level 12 work written by someone else. The credential has to be earned in the role the model will play, level by level.

Every judge's severity estimates cross zero, and most of that is noise

Severity is how harsh or lenient a judge's ratings are at a level, relative to the reference. For every one of the five judges, the severity point estimates sit above zero at some levels and below it at others between levels 5 and 12. On our scale a positive severity means harsher and a negative one means more lenient. Jev's estimates, for example, run from +0.43 logits at level 5 to -0.77 at level 12, and Jev is admitted at neither of those levels; the figure shows each estimate with its error bar.

Five small panels, one per judge (Jev, gpt-oss-120b, qwen3-235b, deepseek-v4.1-flash, and Kimi-k3), each plotting severity in logits for MHC levels 5 to 12 as dots with 95 percent error bars. Bars use the more cautious of two standard errors. Filled dots are admitted levels, hollow dots are excluded ones. Most bars overlap zero; only a few sit clear of it. For example, Jev runs from +0.43 at level 5 to -0.77 at level 12, neither an admitted level. Early evidence, not yet certified.

Now the honest part. Inside the levels where each judge is admitted, no pair of one judge's estimates on opposite sides of zero is further apart than its error bar, once we use the more cautious standard error. A few single levels sit more than their error bar from zero on their own: one judge's estimate at level 9 is on the lenient side by more than its error bar, yet that level is not clearly different from its levels 8 or 10, and no single-level result survives an allowance for the number of cells we checked. Most of the crossings sit inside their own error bars, and every harsh-to-lenient contrast that is larger than its error bar involves at least one level where that judge is not admitted.

That is the lesson, not a letdown. A severity number printed without a standard error invites you to correct noise, and a correction built on noise moves the error around instead of removing it. If a real sign change does exist, one global offset cannot fix it either, because it makes one end worse whichever way you tune it. Either way you need the error bar and the level before you act on a severity number. We have our own gap here. The standard errors behind these cells exist in our earlier credential record, and the figure's error bars use the more cautious standard error for each cell, but they did not carry over when those cells moved into our new database, and our newest runs have not stored them yet. We are fixing both.

What this means for you

If a model scores work that feeds a release, a vendor choice or a promotion decision, check these before you trust the number:

  • Ask for fit and severity at each level of difficulty you judge, never one average.
  • Ask for the standard error on every severity, and treat an estimate inside its error bar as noise.
  • Ask where the judge is not qualified. Holes matter as much as coverage.
  • Treat a fit reading far below 1.0 as a warning about discrimination, not a badge of consistency.
  • Do not infer a judging credential from a model's price or from how well it reasons.
  • Re-check each level after a provider update, because a pooled figure can hide a shift at one level.

Our own status, stated once. Five judges hold credentials under our earlier procedure; re-credentialing under our stricter new procedure is under way. Choosing the least expensive judge that is credentialed at each level, and escalating to a stronger judge only when the evidence requires it, is on our roadmap. Today it is a design, not a running service.

The three-figure self-test

Before your next evaluation cycle, ask whoever supplies your judge for three figures. First, the judge's fit at each level you use it. Second, its severity at each of those levels, with a standard error. Third, the list of levels where it was measured and did not qualify.

If you get all three, you can decide where to trust the judge. If you get one number, you are trusting an average that may be true at no level you care about.

If you want to walk through your own judges against those three figures, book 20 minutes.

This is early evidence: every figure in this post comes from calibrations dated 2026-09-20 (Jev) and 2026-09-23 (the other four judges), and the credentials behind them are being redone under a stricter procedure.

Frequently asked questions

Why is one quality score not enough for an LLM judge?

A judge's fit, and whether it qualifies at all, change from one level of reasoning complexity to the next. In our pilot, one judge fit the measurement model as expected at one level and was far too predictable at the level just below it. An average across levels can describe a judge that exists at no level, so a judge should be qualified level by level.

Does a more expensive or more capable model make a better judge?

Not reliably. In our pilot of five judges, the most expensive, Kimi-k3 (Moonshot AI), was not admitted at a level where three cheaper judges were, and it answered at the top of our earlier reasoning ruler yet was not qualified as a judge at four of the eight levels we ran. Price and reasoning ability do not tell you where a model can judge.

What should I ask a judge vendor for?

Three figures for each level you judge: the judge's fit, its severity with a standard error, and the list of levels where it was measured and did not qualify. A severity estimate that sits inside its standard error should be treated as noise, not corrected.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.