XLNCXLNC Watch it measure

NIST AI 800-3

“Raw benchmark scores do not have this property.”

NIST AI 800-3, February 2026. Quoted verbatim.

What NIST AI 800-3 recommends.

AIM implements at measurement grade the latent-trait/GLMM paradigm NIST AI 800-3 recommends. AIM is not a NIST standard. It claims no endorsement.

What the sentence asks for.

NIST AI 800-3 asks that the distance between two models' capabilities on the scale keep its meaning whichever tasks were picked for the benchmark. On an interval scale, a two-point gap means the same amount everywhere. That is what the sentence asks for. It is what our gate checks.

A worked example, not measured data. The raw score gap moves with the task mix. The gap on the scale stays put.

Worked example, computed from the Rasch model, not measured data. Models A and B sit 2 logits apart on the scale, in every task mix. Raw percent correct: easy mix A 95.3 and B 99.3, 4.0 points apart; middle mix A 50.0 and B 88.1, 38.1 points apart; hard mix A 4.7 and B 26.9, 22.2 points apart.

Where each part sits.

  • A shared scale for questions and models: the Rasch model puts both on one scale.
  • Distances that mean the same thing everywhere: the scale's unit is the logit.
  • A stated error for each result: every measure carries its own error bar.
  • Results that do not hang on one task mix: the questions are locked, and the mix is recorded.
Their recommendation, our implementation. Four parts of the NIST AI 800-3 sentence, each followed by where it sits in AIM: A shared scale for questions and models: The Rasch model puts both on one scale; Distances that mean the same thing everywhere: The scale's unit is the logit; A stated error for each result: Every measure carries its own error bar; Results that do not hang on one task mix: The questions are locked, and the mix is recorded. A property the report describes. NIST has not reviewed, tested or endorsed AIM.

What alignment is not.

AIM is not a NIST standard. NIST has not reviewed, tested or endorsed it. The sentence above describes a property. It does not describe our product.

Two separate records on one time line, 2024 to 2027. NIST lane: February 2026, NIST AI 800-3. AIM lane: 2024, the method published as Chapter 3 of a De Gruyter book (year shown as a span); 2026-08-22, the reasoning ruler's precision gate passed. No review, test or endorsement links the lanes.

The dated record.

Measured on the reasoning ruler. Standard error 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22. Newer scales must clear 0.10 at each level, and 0.15 at level 1.

The method is published as Chapter 3 of a book on models and metrology, edited by Fisher and Pendrill. De Gruyter published it in 2024.

Read the chapter (open access)

Forward this to the person who signs.

The precision gates, drawn to scale as plus or minus one standard error, in levels: 0.07 of a level or better on the reasoning ruler at levels 8 to 12, gate passed 2026-08-22; 0.10 for newer scales at each level; 0.15 for newer scales at level 1. These are gate limits, not individual results.
  1. Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.