Recorded session. 139 questions.
Watch the error bar close.
A recorded adaptive session, start to finish. The ruler picks each question, the estimate updates, the band narrows. Nothing is cut.
The session, start to finish.
The ruler asks. The performer answers. The band narrows. After 139 questions the session stops, because its stop rule is met: level 12 is confirmed.
The performer in this recording is a scripted test performer, not a named model. It shows the engine at work, not a model's result.
Recorded session, final read
Performer: a scripted test performer. Stop rule met: level 12 confirmed. Recorded 2026-08-27.
Trace
Bank fingerprint (sha256, first 16) af70aa616e4a7b66, 4,212 questions in the bank. First question cognitive_complexity_v1_s05.
What you get.
For an admitted judge, the same answer lands at the same point on the ruler, within its error bar. This rests on the judge correction, which has not yet passed its gate.
The correction, in the whitepaper
Measured on the reasoning ruler. Standard error 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22. Newer scales must clear 0.10 at each level, and 0.15 at level 1.
Reasoning ruler, error bar at each level
| Level | Standard error | Judgments |
|---|---|---|
| Level 8 (Concrete) | ± 0.070 | 342 |
| Level 9 (Abstract) | ± 0.056 | 230 |
| Level 10 (Pre-Formal / Abstract-Formal transition) | ± 0.056 | 230 |
| Level 11 (Formal) | ± 0.032 | 232 |
| Level 12 (Systematic) | ± 0.019 | 225 |
No total row. Gate passed 2026-08-22. Unit: one level of the ruler.
Trace
Each row resolves to its question IDs, their answer keys and the bank fingerprint, in the dated report of the certifying run.
Know the cheapest judge that can.
Which judge, at what cost, and when to escalate. Each step in the trace says why that judge and why that question.
Today you pick the judge and see its ceiling. The hand-off to a stronger judge is in build.
In build.
Every dimension, its own error bar.
Each thing measured keeps its own error bar on the same ruler. No total sits on top, so a weak dimension stays in view.
The picker shows the number of questions a run can actually use. Not the number in the bank.
The profile view is in build. Its first real report is not published yet.
In build.
Profile, one error bar per dimension
| Dimension A | not yet published |
| Dimension B | not yet published |
| Dimension C | not yet published |
In build. No total row, by rule.
Your own text, your own number.
Live measurement on your own text is by invitation today. You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.
You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.
What has passed its gate today.
The reasoning ruler has passed its gate. Two more scales are still in calibration: one for hallucination, one for persuasion. Both say so on every page.
If your domain is not on the list, we say so before the run.
Recorded sessions.
More recorded sessions, with their full paths, are open to anyone.
- Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.