XLNCXLNC Watch it measure

MEASUREMENT SCALE 1

Canonical Reasoning Complexity (MHC Stage)

Gate passed 2026-08-22

What it measures. When you are about to trust a model with work that has real structure, you need to know how much complexity it can actually sustain, not how it did on a quiz. This scale measures the developmental order of hierarchical complexity a performer, human or AI, can sustain: the flagship ruler every other AIM scale re-anchors onto. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer. A benchmark can tell you two models tied on a test form. This ruler tells you where each one sits on the complexity order, with the standard error stated at every band. Frozen bank, audit-ready.

All measurement scales · How it works · The Science

S1. Construct definition

Boundary. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer.

Decision context. Feeds capability sufficiency, model selection, and work redesign decisions (the three-decisions block on the Science page).

S2. The Rasch/PCM ruler

Rasch family measurement model; item bank calibrated on real responses, program pre-registered on OSF, stage boundaries anchored to human data.

Measurement frame. Units are logits on the MHC complexity order, with stage-boundary pins set from a published human-data reference (Dawson-Tunik, Commons, Wilson and Fischer, 2005). A one-logit difference is a constant odds ratio everywhere on the ruler: intervals mean the same amount at every stage. The model estimates in logits, and every location and precision figure this page reports, the 0.07 gate included, is stated in levels of the ruler.

Level 8Level 9Level 10Level 11Level 12Levels 8 to 12, with the stage boundaries between them
Gate passed 2026-08-22. MHC Reasoning Complexity - Logic-1 hierarchical. Canonical construct map with stage-boundary pins from the published human-data reference. The precision comparison lives on How it works.

S3. Stated SE per band

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.

BandBand labelStandard errorCertification implication
Bands 8 to 12Gated measurement-method standardSE 0.07 of a level or betterGate passed 2026-08-22. Criterion met: SE at or below 0.07 of a level, with at least 30 judgments, at every band from 8 to 12. Newer scales must clear SE 0.10 per band, band 1 at SE 0.15.
--BELOW FLOOR Per-band judge-evidence readout Model rows whose band severity SE exceeds the frozen quality floor render on the Science page select-ruler, flagged, never presented as a pick.--Shown and flagged, never removed.
--NO EVIDENCE Per-band no-evidence models A model with no per-band judge evidence at a band is listed under no evidence, never silently dropped. Both row classes render on the Science page select-ruler readout.--Shown and flagged, never removed.

S4. Certification tier

Gate passed 2026-08-22

Gate passed 2026-08-22: frozen bank. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.

Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.

S5. Honest-uncertainty display

Rows that fail the floor or lack evidence are shown and flagged, never removed.

The flagged rows below are fixtures standing in for the live select-ruler readout; the rendered rows with real model names are on the Science page.

The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.

S6. API and usage detail

Status. Live in the demo.

POST a response set against the MHC bank; receive a band location on the complexity order with the standard error stated at the value.

POST /score  {"scale": "mhc", "responses": [...]}
{"band": "S11", "location_level": "<stated in levels of the ruler>", "se": "<stated at the value>"}

Rate and plan per the pricing page. The newer scales must clear SE 0.10 per band.

The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.

S7. Scientific references

  1. 2026 Barney, M., Wind, S., & Krishna, V. (2026). Using large language models to evaluate ethical persuasion text: A measurement modeling approach. IJATE, 13(1), 224-247. https://doi.org/10.21449/ijate.1788563
  2. 2016 Barney, M.F. & Fisher, W. P., Jr. (2016). Adaptive Measurement and Assessment. Annual Review of Organizational Psychology and Organizational Behavior, 3, 469-490. https://doi.org/10.1146/annurev-orgpsych-041015-062329
  3. OSF Calibration study pre-registration and audit artifacts (Open Science Framework), per the Science page research program.