Twelve judges. Run underway.
Measure the judge before you trust it.
Every judge we have measured has a ceiling and its own harshness. AIM measures both, so you can start with the cheapest judge that can do the job.
Every judge has a ceiling.
Each judge gets a ceiling, level by level. Two ceilings are published, from judges no longer in our roster. The record stays on the Evidence page. Ceiling rule met
The current judges are in the certified set for levels 7 to 9. Their credentialing run is in progress, so no ceiling is shown for them yet.
The bottom of the range is not mapped yet. The board is regenerated after the next judge runs.
Harsh here, lenient there.
Most judges sit near neutral at most levels. Where a judge is harsh at one level and lenient at another, AIM measures that at each level and removes it there. One whole-model correction would get one end wrong.
In a simulation on our own judge telemetry, harshness left in made a run miss its stated error bar by a wide margin. Corrected, it held.
Jev1 sign change beyond the band
openai-gpt-oss-120b1 sign change beyond the band
qwen3-235b-a22b-instruct-2507no sign change beyond the band
deepseek-v4-1-flashno sign change beyond the band
kimi-k3no sign change beyond the band
Above 0 = harsher. Below 0 = more lenient. Pattern only; the exact values stay in the dated report. Shaded band: plus or minus 2 standard errors. Solid dot: admitted at that level; hollow: not admitted. A ring marks a sign change that clears the band on both sides. Level along the bottom. Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.
Some bias can be fixed. Some must go.
Nine readings of one face. Five are distorted in a consistent way, so each has an exact inverse and can be undone. Three have lost information. No correction brings it back, so they are left out.
Judges work the same way. A judge that leans the same way each time is measured and corrected, level by level.
A judge that rates without a pattern at a level is not used there. Nor is one too vague to place the work.
Correcting harshness alone keeps every judge and every question. AIM also checks fit and precision. What fails is dropped.









Illustration, an analogy, not data. The numbers below the picture are measured. Correctable distortions have an exact inverse; the others do not. Painting: Leonardo da Vinci, Mona Lisa, public-domain scan (C2RMF) via Wikimedia Commons.



Average all eight readings and every distortion survives, smeared into the result. Pixel error is the mean absolute difference from the true image on a 0 to 255 scale, computed from the tiles above. Analogy, not data.
| Kind of bias | What it looks like | What AIM does |
|---|---|---|
| Uniform harshness or leniency | The judge is consistently harsher or softer | Measures it and removes it |
| Harshness that changes with level, sign flips | Harsh at one level, lenient at another | Measures and removes it at each level |
| Inconsistent rating within a level, or misfit | Ratings do not follow how hard the work is | Does not use the judge at that level |
| An error bar too wide at a level | Too blurred to place the work | Does not use the judge at that level |
| Questions that do not fit the ruler | The question does not behave like the others | Rejects the question at the gate |
- Corrected, level by level: Jev is harsher at level 6 and more lenient at level 12. One whole-model correction would get one end wrong.
- Excluded: openai-gpt-oss-120b is admitted at levels 7 and 9 only. At levels 6, 8 and 10 it rates inconsistently within the level, and at level 11 its infit is 2.02, above the critical value. It is never routed there.
- Excluded for precision: no judge on this panel is used at levels 11 or 12. Every error bar there is wider than the target. kimi-k3 at level 11 also underfits (infit 1.62).
Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.
Escalate only when the error bar stays wide.
The rule: the cheapest judge whose ceiling clears the question takes it. A stronger judge is called only when the error bar crosses a level. Status: ceilings are measured today. Automatic routing is not released.
Hand-off runs on the error bar, not on information gain. The next question is chosen for what it teaches. Each step says why.
Cheapest judge first. A stronger judge on stand-by, in build. Stops you set: the widest error bar, the most questions, the most spend.
On a set of 174 located questions, one judge leaves 30 above its ceiling. Paired to judges whose ceilings clear them, all 174 land inside a span.
In build. Proposed.
Your judge, on the ruler.
Send one model and pick one scale. You get its ceiling, its harshness and a dated report.
Judge calibration, one model and one scale: from $2,500.
Tell us which judge you are not sure about.
You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.
The cheapest capable judge.
The One Coarse Judge tier uses the cheapest capable judge, with no escalation. It grades reliably inside the levels it is qualified for, six to ten. Above ten, its scores are not counted.
We measure it the same way as the others. It is admitted today, with a condition.
| Judge | Per 1,336 ratings | L5 | L6 | L7 | L8 | L9 | L10 | L11 | L12 |
|---|---|---|---|---|---|---|---|---|---|
| Jev | US$0.00 | not admitted | admitted, whole level only | admitted, whole level only | admitted | admitted, whole level only | admitted, whole level only | not admitted | not admitted |
| openai-gpt-oss-120b | US$0.50 | not admitted | not admitted | admitted, whole level only | not admitted | admitted, whole level only | not admitted | not admitted | not admitted |
| qwen3-235b-a22b-instruct-2507 | US$1.16 | not admitted | admitted | admitted, whole level only | admitted, whole level only | admitted | admitted, whole level only | not admitted | not admitted |
| deepseek-v4-1-flash | US$4.41 | not admitted | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted | not admitted | not admitted |
| kimi-k3 | US$29.75 | not admitted | not admitted | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted, whole level only | not admitted | not admitted |
No shared blind spot.
The judges come from many vendors in several countries. The companies that write and attack the new questions do not judge them. No judge scores a question its own company wrote. So no one vendor's blind spot or self-preference decides a score.
The questions are fresh. We write new ones automatically, after the models were trained. We keep them sealed and never publish them. So they could not have been memorized.
Countries are the vendors' home countries.
Judges, levels 7 to 9
In the certified set. Credentialing run in progress, no results yet.
- Gemma 4 31B Google
- Gemma 4 26B Google
- Gemma 3 27B Google
- GPT-6 Luna
OpenAI
- GPT-OSS 120B
OpenAI
- Mercury 2.5 Inception
- Mistral Small 3.2
Mistral
- Nemotron 3 Nano 30B NVIDIA
- MiMo V2.5 Xiaomi
- Qwen 3 8 Flash Alibaba level 8 only
- Qwen3 235B Alibaba level 8 only
- Jev TypeSafe
Write and attack the questions
- Writes and repairs the questions
Anthropic
- Attacks the questions
xAI if it passes a reliability check; Anthropic as fallback
Judges, levels 10 to 13
Candidates being screened. Not a final roster.
- A subset of the level 7 to 9 judges levels 10 to 12
- DeepSeek V4.1 Flash DeepSeek levels 10 to 13
- GLM 5.3 Flash Z.ai levels 10 to 12
- Kimi K3
Moonshot level 12 and up
- GPT-6 Astra
OpenAI level 13, one-level trial first
Roster as of 2026-10-02. No Anthropic model judges at any level. No xAI model judges at levels 10 to 13. Flags show each vendor's home country.
Vendor and model names are used only to identify the models XLNC measures. XLNC is not affiliated with, sponsored by, or endorsed by these companies, and their inclusion does not imply their approval of any result. All names belong to their respective owners.
- Ceiling rule met. Criterion: a judge's ceiling is the highest level it passes before a break; passes above a break do not count. Certifying runs: the reasoning-ceiling probe runs of 2026-08-19 and 2026-08-22, audited by our quality gate.
