XLNCXLNC Watch it measure

Gate passed 2026-08-22

Modelfree to move
Judgemeasured
Questionslocked, with a fingerprint
139questions
12.40measured level
0.153error bar (one standard error)

Band: a recorded session on a scripted test performer, 139 questions, 2026-08-27.

Your score moved. What moved it?

AIM measures the judge, locks the questions and puts an error bar on each number. A move bigger than its error bar points at the model.

Today this works on our own locked question sets. Your own judge can be measured the same way, so its scores come with an error bar too.

Some bias can be fixed. Some must go.

A judge that leans the same way each time is corrected, level by level. A judge that rates without a pattern at a level is not used there.

See how, on the Judges page

The Mona Lisa, undistorted: the calibrated measure
True imageReferenceThe calibrated measure
The Mona Lisa, stretched: stands for uniform leniency; correctable
StretchedCorrectableUniform leniency
The Mona Lisa, warped as in a funhouse mirror: stands for hallucination; exclude, cannot be corrected
HallucinationExclude: cannot be correctedFunhouse mirror

Every judge bends the picture. AIM undoes the bends it can measure and stops using a judge where it cannot. Illustration, an analogy. Painting: Leonardo da Vinci, Mona Lisa, public-domain scan (C2RMF) via Wikimedia Commons.

Judge measured. Questions locked.

Four things can move a score: the model, the judge, the questions or the prompt. Here the judge is measured and the questions and their keys are locked. Chance shows up as the error bar. What is left points at the model.

  • Model: free to move.
  • Judge: measured.
  • Questions and keys: locked, with a fingerprint.
  • Grading prompt: held constant. Its record is pending.

Your own task mix in production is outside this ruler.

In a simulation on our own judge telemetry, a run set to a standard error of 0.20 came in at 1.16 with judge harshness left in. With it corrected, 0.21.1

Simulation.

789101112131415 harsher judgemore lenient judge After correction both readings meet. The ruler does not move.

Two questions to ask any ruler.

Has the ruler moved?

The questions cannot change.

The question set is frozen, fingerprinted and dated. Every run checks the fingerprint first. Whether the ruler holds steady over months is the next exhibit, and it is not done yet.

Re-measurement run: not yet published.2

Is the ruler circular?

Defined before anything is scored.

Each level is fixed before any answer is graded. New questions stay sealed, so no model could have learned them. For new questions, the writer, the attacker and the judges come from separate companies. Each judge is checked against the others.

The judges for levels 7 to 9 are a certified set of 12. Their credentialing run is in progress, with no results yet.

A check against outside human experts: still ahead.

See the roster

Measured, not averaged.

Each level carries its own error bar. No average sits on top. A failure at one level stays visible.

Pick any score. It resolves to one question in a locked question set, with its answer key and a fingerprint.

Measured on the reasoning ruler. Gate passed 2026-08-22. Each level's error bar is on the Evidence page, with its unit.

If your domain is not on the list, we say so before the run.

Reasoning ruler, error bar at each level

LevelStandard errorJudgments
Level 8 (Concrete)± 0.070342
Level 9 (Abstract)± 0.056230
Level 10 (Pre-Formal / Abstract-Formal transition)± 0.056230
Level 11 (Formal)± 0.032232
Level 12 (Systematic)± 0.019225

No total row. Gate passed 2026-08-22. Unit: one level of the ruler.

Trace

Each row resolves to its question IDs, their answer keys and the bank fingerprint, in the dated report of the certifying run.

Know the cheapest judge that can.

Two judges have published ceilings. Today you pick the judge and see how far up the ruler it grades reliably. The run stops when the error bar is tight enough, or at the question limit you set. Hand-off to a stronger judge is in build.

Cheapest judge first. A stronger judge on stand-by, in build. Stops you set: the widest error bar, the most questions, the most spend.3

In build.

cheapest capable judge stronger judge strongest judge stop: widest error barstop: most questionsstop: most spend In build

Open to reviewers. Closed to training.

The engine is open to reviewers under Apache-2.0, so reviewers can check the method. The banks stay closed. That keeps the questions out of public training data.

If you run an eval loop, AIM is one step in it. It replaces nothing you built.

The whole loop, one click down.

Engine

Apache-2.0
open to reviewers

Question banks

closed
sha256 fingerprint

One outside view.

Headshot of the person quoted.
“I've always stressed the importance of ethics in persuasion, and Dr. Matt Barney's AI assessment tool brings unprecedented scientific rigor to this domain. I am optimistic that his method holds immense promise in proactively preventing the misuse of persuasion techniques, both by people and emerging technologies, and augmenting their long-term use correctly”
Dr. Robert Cialdini, NY Times Best Selling Author, Founder Cialdini Institute, Regents Professor Emeritus ASU

Every measure, with its date.

Every measure on this page carries its error bar and its date. When a number moves, you can see which parts were held still.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

Forward this to the person who signs.

  1. Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.
  2. Ceiling rule met. Criterion: a judge's ceiling is the highest level it passes before a break; passes above a break do not count. Certifying runs: the reasoning-ceiling probe runs of 2026-08-19 and 2026-08-22, audited by our quality gate.
  3. The simulation used 15,000 judge decisions, dated 2026-09-09. Method and results are in the whitepaper. Whitepaper
  4. Fifteen artifacts sit under sha256 fingerprints, checked at the start of every run.
  5. Two ceilings are published. Both sit at level 11, with a stated condition. Ceiling rule met
  6. Method: Chapter 3 of a book on models and metrology. Edited by Fisher and Pendrill. Published by De Gruyter in 2024.