Two judges. One ruler.
A ruler that stretches is not a ruler.
Every judge we have measured is harsh in some places and lenient in others. AIM measures that pattern first and corrects for it.
Five problems with AI scores today.
- No error barA score, but is it reliable or comparable?1Traceable measurementEvery published result carries its standard error.
- Averages hide failureA healthy mean masks the level that breaks2Measured by levelResults located on the MHC stage ruler.
- Capability is jaggedStrong at some levels, missing at others3Credentialed by levelEach judge used only where it is credentialed.
- Judges drift and disagreeHarsh at one level, lenient at the next4Calibrated AI-as-JudgeEach judge's harshness measured and removed, level by level.
- Tests leak and get gamedA static benchmark decays on release5Fresh questionsNew questions placed on the ruler before first use.
Equal steps, everywhere.
A two-point gap means the same amount everywhere on the ruler. A score cannot promise that. A measure must.
The test adapts to the model.
The test picks the next question from what it has learned so far. It stops when the answer is precise enough.
A short clip of the test choosing its next question is still to come.
A recorded session, 139 questions.
Writers never grade.
The model that writes a question never grades it. The graders come from other model families.
No judge scores a question its own company wrote. So no one vendor's blind spot decides the score.
Validity study: not yet shown.
Judges, levels 7 to 9
In the certified set. Credentialing run in progress, no results yet.
- Gemma 4 31B Google
- Gemma 4 26B Google
- Gemma 3 27B Google
- GPT-6 Luna
OpenAI
- GPT-OSS 120B
OpenAI
- Mercury 2.5 Inception
- Mistral Small 3.2
Mistral
- Nemotron 3 Nano 30B NVIDIA
- MiMo V2.5 Xiaomi
- Qwen 3 8 Flash Alibaba level 8 only
- Qwen3 235B Alibaba level 8 only
- Jev TypeSafe
Write and attack the questions
- Writes and repairs the questions
Anthropic
- Attacks the questions
xAI if it passes a reliability check; Anthropic as fallback
Judges, levels 10 to 13
Candidates being screened. Not a final roster.
- A subset of the level 7 to 9 judges levels 10 to 12
- DeepSeek V4.1 Flash DeepSeek levels 10 to 13
- GLM 5.3 Flash Z.ai levels 10 to 12
- Kimi K3
Moonshot level 12 and up
- GPT-6 Astra
OpenAI level 13, one-level trial first
Roster as of 2026-10-02. No Anthropic model judges at any level. No xAI model judges at levels 10 to 13. Flags show each vendor's home country.
Vendor and model names are used only to identify the models XLNC measures. XLNC is not affiliated with, sponsored by, or endorsed by these companies, and their inclusion does not imply their approval of any result. All names belong to their respective owners.
Every score traces back.
Pick any score. It resolves to one question in a locked question set, with its answer key and a fingerprint.
Recorded session, final read
Performer: a scripted test performer. Stop rule met: level 12 confirmed. Recorded 2026-08-27.
Trace
Bank fingerprint (sha256, first 16) af70aa616e4a7b66, 4,212 questions in the bank. First question cognitive_complexity_v1_s05.
Your loop. Our ruler.
Trigger, design, build, measure, decide, monitor, iterate. AIM is the measure step. If you run an eval loop, AIM is one step in it. It replaces nothing you built.
- Trigger: a new model, a new judge or a new release asks for a number.
- Design: pick the scale and the stops you will accept.
- Build: point the run at the model and the judge.
- Measure: the test adapts until the error bar is tight enough.
- Decide: is the model good enough for this job, and which model to pick.
- Monitor: measure again when something changes.
- Iterate: change one part and read which part moved.
A headless API and MCP endpoint is planned. It is not callable today. Do not write a tool call against it today.
Select, upstream
Adaptive, overt test
- Before deployment: choose or qualify a model.
- The model answers calibrated questions, each chosen for what it teaches.
- Credentialed judges score; each judge's harshness is removed.
“Should we ship this model?”
Operate, downstream In build
Passive measurement of the work
- In production: track capability and drift, window by window.
- No test questions injected: it measures the outputs the system already produces.
- Passive, inverted CAT: the method picks which real outputs to score next, the way an adaptive test picks questions.
“Is it still performing?”
Select with the test. Operate with the work it already produces. Same ruler, a measure plus its error bar.
What the ruler measures.
- Canonical Reasoning Complexity (MHC Stage)Gate passed 2026-08-22
- Hallucination ScaleIn calibration
- Persuasion BatteryIn calibration
- Agent-Generated Research Certification (AGRC)Watch
- Rule Retention Coefficient (RRC)Watch
- Agent Eval / Agentic-Loop ReliabilityWatch
Gate passed means a stated rule was met in a named run, and a quality check signed it. In calibration means the numbers are real and the gate run has not yet passed. Watch means the skill is defined and not yet measured.
If your domain is not on the list, we say so before the run.
Beside other instruments.
How tight a measurement is, compared with the decision it supports, is one axis. On that axis this ruler sits beside physical gauges. Sources are listed below. This is not a claim that AIM beats any named test.
AIM is not a NIST or ISO standard. The claim is only that the error is stated with the same care. We are not tied to, backed by or paid by any group named here.
Physical figures: ISO 3650, ISO 3611, ISO 13385-1, and the SI metre as realized under the 1983 definition. The ruler: 0.07 of a level, from the certifying run.
Tell us which score you could not explain.
You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.
- Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.