XLNCXLNC Watch it measure

Two judges. One ruler.

789101112131415 harsher judgemore lenient judge After correction both readings meet. The ruler does not move.

A ruler that stretches is not a ruler.

Every judge we have measured is harsh in some places and lenient in others. AIM measures that pattern first and corrects for it.

Five problems with AI scores today.

  1. No error barA score, but is it reliable or comparable?
    1Traceable measurementEvery published result carries its standard error.
  2. Averages hide failureA healthy mean masks the level that breaks
    2Measured by levelResults located on the MHC stage ruler.
  3. Capability is jaggedStrong at some levels, missing at others
    3Credentialed by levelEach judge used only where it is credentialed.
  4. Judges drift and disagreeHarsh at one level, lenient at the next
    4Calibrated AI-as-JudgeEach judge's harshness measured and removed, level by level.
  5. Tests leak and get gamedA static benchmark decays on release
    5Fresh questionsNew questions placed on the ruler before first use.
Why AI scores are hard to trust today, and how AIM addresses each one. Measurement and evaluation, 24 September 2026.

Equal steps, everywhere.

A two-point gap means the same amount everywhere on the ruler. A score cannot promise that. A measure must.

How the model does it

a measure: even steps a raw score: uneven steps

The test adapts to the model.

The test picks the next question from what it has learned so far. It stops when the answer is precise enough.

A short clip of the test choosing its next question is still to come.

139questions
12.40measured level
0.153error bar (one standard error)

A recorded session, 139 questions.

Writers never grade.

The model that writes a question never grades it. The graders come from other model families.

No judge scores a question its own company wrote. So no one vendor's blind spot decides the score.

Validity study: not yet shown.

Judges, levels 7 to 9

In the certified set. Credentialing run in progress, no results yet.

  • Gemma 4 31B Google
  • Gemma 4 26B Google
  • Gemma 3 27B Google
  • GPT-6 Luna OpenAI
  • GPT-OSS 120B OpenAI
  • Mercury 2.5 Inception
  • Mistral Small 3.2 Mistral
  • Nemotron 3 Nano 30B NVIDIA
  • MiMo V2.5 Xiaomi
  • Qwen 3 8 Flash Alibaba level 8 only
  • Qwen3 235B Alibaba level 8 only
  • Jev TypeSafe

Write and attack the questions

  • Writes and repairs the questions Anthropic
  • Attacks the questions xAI if it passes a reliability check; Anthropic as fallback

Judges, levels 10 to 13

Candidates being screened. Not a final roster.

  • A subset of the level 7 to 9 judges levels 10 to 12
  • DeepSeek V4.1 Flash DeepSeek levels 10 to 13
  • GLM 5.3 Flash Z.ai levels 10 to 12
  • Kimi K3 Moonshot level 12 and up
  • GPT-6 Astra OpenAI level 13, one-level trial first

Roster as of 2026-10-02. No Anthropic model judges at any level. No xAI model judges at levels 10 to 13. Flags show each vendor's home country.

Vendor and model names are used only to identify the models XLNC measures. XLNC is not affiliated with, sponsored by, or endorsed by these companies, and their inclusion does not imply their approval of any result. All names belong to their respective owners.

Every score traces back.

Pick any score. It resolves to one question in a locked question set, with its answer key and a fingerprint.

The trace chain, on The Science

Recorded session, final read

12.40measured level
± 0.153error bar
139questions

Performer: a scripted test performer. Stop rule met: level 12 confirmed. Recorded 2026-08-27.

Trace

Bank fingerprint (sha256, first 16) af70aa616e4a7b66, 4,212 questions in the bank. First question cognitive_complexity_v1_s05.

Your loop. Our ruler.

Trigger, design, build, measure, decide, monitor, iterate. AIM is the measure step. If you run an eval loop, AIM is one step in it. It replaces nothing you built.

  • Trigger: a new model, a new judge or a new release asks for a number.
  • Design: pick the scale and the stops you will accept.
  • Build: point the run at the model and the judge.
  • Measure: the test adapts until the error bar is tight enough.
  • Decide: is the model good enough for this job, and which model to pick.
  • Monitor: measure again when something changes.
  • Iterate: change one part and read which part moved.

A headless API and MCP endpoint is planned. It is not callable today. Do not write a tool call against it today.

Select, upstream

Adaptive, overt test

  • Before deployment: choose or qualify a model.
  • The model answers calibrated questions, each chosen for what it teaches.
  • Credentialed judges score; each judge's harshness is removed.

“Should we ship this model?”

Operate, downstream In build

Passive measurement of the work

  • In production: track capability and drift, window by window.
  • No test questions injected: it measures the outputs the system already produces.
  • Passive, inverted CAT: the method picks which real outputs to score next, the way an adaptive test picks questions.

“Is it still performing?”

Select with the test. Operate with the work it already produces. Same ruler, a measure plus its error bar.

What the ruler measures.

Gate passed means a stated rule was met in a named run, and a quality check signed it. In calibration means the numbers are real and the gate run has not yet passed. Watch means the skill is defined and not yet measured.

If your domain is not on the list, we say so before the run.

Beside other instruments.

How tight a measurement is, compared with the decision it supports, is one axis. On that axis this ruler sits beside physical gauges. Sources are listed below. This is not a claim that AIM beats any named test.

AIM is not a NIST or ISO standard. The claim is only that the error is stated with the same care. We are not tied to, backed by or paid by any group named here.

SI metreGauge block, grade 0MicrometerVernier caliperthe reasoning ruler10:1 rule4:1 ruletighterlooser

Physical figures: ISO 3650, ISO 3611, ISO 13385-1, and the SI metre as realized under the 1983 definition. The ruler: 0.07 of a level, from the certifying run.

Tell us which score you could not explain.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

Forward this to the person who signs.

  1. Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.