One page
Reasoning ruler
Hallucination and persuasion scales: in calibration. Judge hand-off: in build.
AIM on one page.
What Adaptive Intelligent Measurement evaluates, what it costs, what has passed its gate and what has not yet. If your evaluator sent you this, here is the short version. Every line is on the site with its status.
The question behind an eval bill.
If a score moves and no one can say why, the release waits or ships on a guess. That cost is yours to name. This page shows what AIM can settle today.
What it is.
A score can change because the model changed, the judge changed, or the questions changed. AIM measures the judge and locks the questions. What is left points at the model.
Measured on the reasoning ruler. Standard error 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22. Newer scales must clear 0.10 at each level, and 0.15 at level 1.
What it changes.
If your metrics do not move together, an average can hide the one that fell. Each dimension keeps its own error bar, so the one that fell stays visible.
In build.
What it costs.
If the evaluation bill gets argued, the stops are yours to set. The widest error bar you accept. The most questions. The most spend.
Escalation with these stops is part of the plans from the second hosted tier up. Automatic routing is in build, and it is included when released. Prices are on the Pricing page.
What has passed, and what has not yet.
- The reasoning ruler: Gate passed 2026-08-22.
- Hallucination and persuasion scales: gate not yet passed.
- Judge hand-off and stops: in build.
- Per-dimension profile: in build.
- The same model measured on two dates: not yet shown.
- Ruler stability over months, and a validity study: not yet shown.
What we will not claim.
AIM implements at measurement grade the latent-trait/GLMM paradigm NIST AI 800-3 recommends. AIM is not a NIST standard. It claims no endorsement.
- No customer names yet.
- We do not publish a cost of getting this wrong. We have not measured yours.
- No claim about your own prompts or your own task mix.
What it takes to run.
- Where the data goes. Hosted plans run on our servers. The self-hosted harness keeps your data on your side.
- What your engineer has to run. To be confirmed.
- Support contact and response time. To be confirmed.
What happens next.
Watch a recorded session together. Then talk to us about a proof engagement on your hardest case.
Email [email protected] to set up the call.
- Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.