XLNCXLNC Watch it measure

Twelve judges. Run underway.

Levels 7 to 912judges in the certified set
Levels 7 to 9In progresscredentialing run, no results yet
Levels 10 to 13Screeningcandidate judges

Measure the judge before you trust it.

Every judge we have measured has a ceiling and its own harshness. AIM measures both, so you can start with the cheapest judge that can do the job.

Every judge has a ceiling.

Each judge gets a ceiling, level by level. Two ceilings are published, from judges no longer in our roster. The record stays on the Evidence page. Ceiling rule met

The current judges are in the certified set for levels 7 to 9. Their credentialing run is in progress, so no ceiling is shown for them yet.

The bottom of the range is not mapped yet. The board is regenerated after the next judge runs.

The current roster

The published record

Credential grid, five judges by MHC level 5 to 12. deepseek-v4-1-flash: credentialed 6 to 10. kimi-k3: 7 to 10. gpt-oss-120b: 7 and 9 only, with level 8 between them excluded. qwen3-235b-a22b-instruct: 6 to 10. Jev: 6 to 10. No judge is credentialed at 5, 11 or 12.
Five faster judges, credentialed level by level. Filled: credentialed, so the selector can route to that judge at that level. A dash: excluded, never routed. A gap between two filled cells is a hole, and the selector never routes a judge across it.

Four questions to ask your eval judge.

Is the judge measured? How wide is its error bar? Does it fit at each level? Where can it work?

Four questions to ask your eval judge, and what a measurement system reports for each. 1, a credentialed judge: Jev's severity at levels 5 to 8 with error bars; admitted at 6, 7 and 8, not at 5. 2, an error bar and a floor: before and after bars with the smallest detectable improvement; illustrative, not measured data. 3, fit to the model by level: Jev's infit, 0.13 at level 7, flagged overfit and kept at coarse grain only, and 0.97 at level 8. 4, headroom: one or more judges admitted at levels 6 to 10; none at 5, 11 or 12.
Each panel shows what a measurement system reports. Panel 1: above 0 = harsher. Panel 2 is illustrative, as labeled. Calibrations of 2026-09-20 and 2026-09-23.

Harsh here, lenient there.

Most judges sit near neutral at most levels. Where a judge is harsh at one level and lenient at another, AIM measures that at each level and removes it there. One whole-model correction would get one end wrong.

In a simulation on our own judge telemetry, harshness left in made a run miss its stated error bar by a wide margin. Corrected, it held.

The numbers, on the Evidence page

Jev1 sign change beyond the band

+0−6789101112

openai-gpt-oss-120b1 sign change beyond the band

+0−6789101112

qwen3-235b-a22b-instruct-2507no sign change beyond the band

+0−6789101112

deepseek-v4-1-flashno sign change beyond the band

+0−6789101112

kimi-k3no sign change beyond the band

+0−6789101112

Above 0 = harsher. Below 0 = more lenient. Pattern only; the exact values stay in the dated report. Shaded band: plus or minus 2 standard errors. Solid dot: admitted at that level; hollow: not admitted. A ring marks a sign change that clears the band on both sides. Level along the bottom. Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.

Some bias can be fixed. Some must go.

Nine readings of one face. Five are distorted in a consistent way, so each has an exact inverse and can be undone. Three have lost information. No correction brings it back, so they are left out.

Judges work the same way. A judge that leans the same way each time is measured and corrected, level by level.

A judge that rates without a pattern at a level is not used there. Nor is one too vague to place the work.

Correcting harshness alone keeps every judge and every question. AIM also checks fit and precision. What fails is dropped.

The Mona Lisa, washed out: stands for squeezed scale; correctable
Washed outCorrectableSqueezed scale
The Mona Lisa, upside down: stands for harshness flips by level; correctable
Upside downCorrectableHarshness flips by level
The Mona Lisa, sheared: stands for harshness grows by level; correctable
ShearedCorrectableHarshness grows by level
The Mona Lisa, stretched: stands for uniform leniency; correctable
StretchedCorrectableUniform leniency
The Mona Lisa, undistorted: the calibrated measure
True imageReferenceThe calibrated measure
The Mona Lisa, squashed: stands for uniform harshness; correctable
SquashedCorrectableUniform harshness
The Mona Lisa, blurred: stands for unstable judge; exclude
BlurredExcludeUnstable judge
The Mona Lisa, pixelated: stands for out of its depth; exclude
PixelatedExcludeOut of its depth
The Mona Lisa, warped as in a funhouse mirror: stands for hallucination; exclude, cannot be corrected
HallucinationExclude: cannot be correctedFunhouse mirror

Illustration, an analogy, not data. The numbers below the picture are measured. Correctable distortions have an exact inverse; the others do not. Painting: Leonardo da Vinci, Mona Lisa, public-domain scan (C2RMF) via Wikimedia Commons.

1. As rated: Eight readings, no correction. Pixel error against the true image 27.6.
1. As ratedEight readings, no correctionPixel error: 27.6
2. Harshness corrected: Five inverted, three still in. Pixel error against the true image 8.1.
2. Harshness correctedFive inverted, three still inPixel error: 8.1
3. Three excluded: Five corrected, three dropped. Pixel error against the true image 1.5.
3. Three excludedFive corrected, three droppedPixel error: 1.5

Average all eight readings and every distortion survives, smeared into the result. Pixel error is the mean absolute difference from the true image on a 0 to 255 scale, computed from the tiles above. Analogy, not data.

Kind of biasWhat it looks likeWhat AIM does
Uniform harshness or leniencyThe judge is consistently harsher or softerMeasures it and removes it
Harshness that changes with level, sign flipsHarsh at one level, lenient at anotherMeasures and removes it at each level
Inconsistent rating within a level, or misfitRatings do not follow how hard the work isDoes not use the judge at that level
An error bar too wide at a levelToo blurred to place the workDoes not use the judge at that level
Questions that do not fit the rulerThe question does not behave like the othersRejects the question at the gate
  • Corrected, level by level: Jev is harsher at level 6 and more lenient at level 12. One whole-model correction would get one end wrong.
  • Excluded: openai-gpt-oss-120b is admitted at levels 7 and 9 only. At levels 6, 8 and 10 it rates inconsistently within the level, and at level 11 its infit is 2.02, above the critical value. It is never routed there.
  • Excluded for precision: no judge on this panel is used at levels 11 or 12. Every error bar there is wider than the target. kimi-k3 at level 11 also underfits (infit 1.62).

Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.

Escalate only when the error bar stays wide.

The rule: the cheapest judge whose ceiling clears the question takes it. A stronger judge is called only when the error bar crosses a level. Status: ceilings are measured today. Automatic routing is not released.

Hand-off runs on the error bar, not on information gain. The next question is chosen for what it teaches. Each step says why.

Cheapest judge first. A stronger judge on stand-by, in build. Stops you set: the widest error bar, the most questions, the most spend.

On a set of 174 located questions, one judge leaves 30 above its ceiling. Paired to judges whose ceilings clear them, all 174 land inside a span.

In build. Proposed.

cheapest capable judge stronger judge strongest judge stop: widest error barstop: most questionsstop: most spend In build

Your judge, on the ruler.

Send one model and pick one scale. You get its ceiling, its harshness and a dated report.

Judge calibration, one model and one scale: from $2,500.

Tell us which judge you are not sure about.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

The cheapest capable judge.

The One Coarse Judge tier uses the cheapest capable judge, with no escalation. It grades reliably inside the levels it is qualified for, six to ten. Above ten, its scores are not counted.

We measure it the same way as the others. It is admitted today, with a condition.

One Coarse Judge, on the Pricing page

Price does not tell you where a judge works.
JudgePer 1,336 ratingsL5L6L7L8L9L10L11L12
JevUS$0.00not admittedadmitted, whole level onlyadmitted, whole level onlyadmittedadmitted, whole level onlyadmitted, whole level onlynot admittednot admitted
openai-gpt-oss-120bUS$0.50not admittednot admittedadmitted, whole level onlynot admittedadmitted, whole level onlynot admittednot admittednot admitted
qwen3-235b-a22b-instruct-2507US$1.16not admittedadmittedadmitted, whole level onlyadmitted, whole level onlyadmittedadmitted, whole level onlynot admittednot admitted
deepseek-v4-1-flashUS$4.41not admittedadmitted, whole level onlyadmitted, whole level onlyadmitted, whole level onlyadmitted, whole level onlyadmittednot admittednot admitted
kimi-k3US$29.75not admittednot admittedadmitted, whole level onlyadmitted, whole level onlyadmitted, whole level onlyadmitted, whole level onlynot admittednot admitted
● admitted. ◑ admitted at whole-level grain only. × not admitted. Cost is measured from the run ledger. Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.

No shared blind spot.

The judges come from many vendors in several countries. The companies that write and attack the new questions do not judge them. No judge scores a question its own company wrote. So no one vendor's blind spot or self-preference decides a score.

The questions are fresh. We write new ones automatically, after the models were trained. We keep them sealed and never publish them. So they could not have been memorized.

Countries are the vendors' home countries.

Judges, levels 7 to 9

In the certified set. Credentialing run in progress, no results yet.

  • Gemma 4 31B Google
  • Gemma 4 26B Google
  • Gemma 3 27B Google
  • GPT-6 Luna OpenAI
  • GPT-OSS 120B OpenAI
  • Mercury 2.5 Inception
  • Mistral Small 3.2 Mistral
  • Nemotron 3 Nano 30B NVIDIA
  • MiMo V2.5 Xiaomi
  • Qwen 3 8 Flash Alibaba level 8 only
  • Qwen3 235B Alibaba level 8 only
  • Jev TypeSafe

Write and attack the questions

  • Writes and repairs the questions Anthropic
  • Attacks the questions xAI if it passes a reliability check; Anthropic as fallback

Judges, levels 10 to 13

Candidates being screened. Not a final roster.

  • A subset of the level 7 to 9 judges levels 10 to 12
  • DeepSeek V4.1 Flash DeepSeek levels 10 to 13
  • GLM 5.3 Flash Z.ai levels 10 to 12
  • Kimi K3 Moonshot level 12 and up
  • GPT-6 Astra OpenAI level 13, one-level trial first

Roster as of 2026-10-02. No Anthropic model judges at any level. No xAI model judges at levels 10 to 13. Flags show each vendor's home country.

Vendor and model names are used only to identify the models XLNC measures. XLNC is not affiliated with, sponsored by, or endorsed by these companies, and their inclusion does not imply their approval of any result. All names belong to their respective owners.

  1. Ceiling rule met. Criterion: a judge's ceiling is the highest level it passes before a break; passes above a break do not count. Certifying runs: the reasoning-ceiling probe runs of 2026-08-19 and 2026-08-22, audited by our quality gate.