MEASUREMENT SCALE 1
Gate passed 2026-08-22
What it measures. When you are about to trust a model with work that has real structure, you need to know how much complexity it can actually sustain, not how it did on a quiz. This scale measures the developmental order of hierarchical complexity a performer, human or AI, can sustain: the flagship ruler every other AIM scale re-anchors onto. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer. A benchmark can tell you two models tied on a test form. This ruler tells you where each one sits on the complexity order, with the standard error stated at every band. Frozen bank, audit-ready.
Boundary. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer.
Decision context. Feeds capability sufficiency, model selection, and work redesign decisions (the three-decisions block on the Science page).
Rasch family measurement model; item bank calibrated on real responses, program pre-registered on OSF, stage boundaries anchored to human data.
Measurement frame. Units are logits on the MHC complexity order, with stage-boundary pins set from a published human-data reference (Dawson-Tunik, Commons, Wilson and Fischer, 2005). A one-logit difference is a constant odds ratio everywhere on the ruler: intervals mean the same amount at every stage. The model estimates in logits, and every location and precision figure this page reports, the 0.07 gate included, is stated in levels of the ruler.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| Bands 8 to 12 | Gated measurement-method standard | SE 0.07 of a level or better | Gate passed 2026-08-22. Criterion met: SE at or below 0.07 of a level, with at least 30 judgments, at every band from 8 to 12. Newer scales must clear SE 0.10 per band, band 1 at SE 0.15. |
| -- | BELOW FLOOR Per-band judge-evidence readout Model rows whose band severity SE exceeds the frozen quality floor render on the Science page select-ruler, flagged, never presented as a pick. | -- | Shown and flagged, never removed. |
| -- | NO EVIDENCE Per-band no-evidence models A model with no per-band judge evidence at a band is listed under no evidence, never silently dropped. Both row classes render on the Science page select-ruler readout. | -- | Shown and flagged, never removed. |
Gate passed 2026-08-22
Gate passed 2026-08-22: frozen bank. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged rows below are fixtures standing in for the live select-ruler readout; the rendered rows with real model names are on the Science page.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. Live in the demo.
POST a response set against the MHC bank; receive a band location on the complexity order with the standard error stated at the value.
POST /score {"scale": "mhc", "responses": [...]}
{"band": "S11", "location_level": "<stated in levels of the ruler>", "se": "<stated at the value>"}
Rate and plan per the pricing page. The newer scales must clear SE 0.10 per band.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.