MEASUREMENT AND EVALUATION
One page per scale: what it measures first, then the ruler, the stated standard error per band, the tier and the honest-uncertainty display. Rows that fail the floor or lack evidence are shown and flagged, never removed.
Sorted by status: gate passed first, then in calibration, then watch. No filtering: Watch scales are listed, not hidden. Every number on a scale page carries a tier word or a source; where no gated figure exists, the page says so.
Gate passed 2026-08-22
How much complexity a performer, human or AI, can actually sustain.
Precision: SE 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22
In calibration
How often a model asserts unsupported content, and where it stops.
Precision: best band SE stated on page
1 no-evidence row, shown
In calibration
How hard content pushes, and whether it pushes ethically.
Precision: best band SE stated on page
1 no-evidence row, shown
Watch
Evidence, with its error bar, of what an agent got right in research outputs.
Precision: none certified
1 no-evidence row, shown
Watch
How much of an agent's stated rules survive repeated context compaction.
Precision: none certified
1 no-evidence row, shown
Watch
Whether an agent finishes what it started, in order, within constraints.
Precision: none certified
1 no-evidence row, shown