NIST AI 800-3
“Raw benchmark scores do not have this property.”
NIST AI 800-3, February 2026. Quoted verbatim.
What NIST AI 800-3 recommends.
AIM implements at measurement grade the latent-trait/GLMM paradigm NIST AI 800-3 recommends. AIM is not a NIST standard. It claims no endorsement.
What the sentence asks for.
NIST AI 800-3 asks that the distance between two models' capabilities on the scale keep its meaning whichever tasks were picked for the benchmark. On an interval scale, a two-point gap means the same amount everywhere. That is what the sentence asks for. It is what our gate checks.
A worked example, not measured data. The raw score gap moves with the task mix. The gap on the scale stays put.
Where each part sits.
- A shared scale for questions and models: the Rasch model puts both on one scale.
- Distances that mean the same thing everywhere: the scale's unit is the logit.
- A stated error for each result: every measure carries its own error bar.
- Results that do not hang on one task mix: the questions are locked, and the mix is recorded.
The dated record.
Measured on the reasoning ruler. Standard error 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22. Newer scales must clear 0.10 at each level, and 0.15 at level 1.
The method is published as Chapter 3 of a book on models and metrology, edited by Fisher and Pendrill. De Gruyter published it in 2024.
- Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.



