Adaptive Intelligent Measurement (AIM): questions and answers
Adaptive Intelligent Measurement (AIM) is built to place people and AI systems on one scale; AI systems are measured today, people are not yet, each with a stated error. Answers below state their evidence and limits. The dated gate record is on the Evidence page. One Ruler. Every Mind.
I am a...
AI Evaluation Engineers
- Can the number be trusted?
- Was it the model, the judge, the data, the task mix, or my own test prompt?
- Did my model or agent get better or worse?
- How do you stop a model from gaming or memorizing the questions?
- Is the ruler circular, graded by the same kind of models it measures?
- My outputs are long or subjective. How do you judge them?
- What does a run cost?
- How is this different from Trismik or the eval tools in my platform?
- What is calibrated today, and what is not?
- How does output reach me? Is there an API?
- Who checks the checker, and what happens when a gate fails?
- Can I bring my own judge and keep data inside my boundary?
AI Safety and Governance leads
- Can the number be trusted?
- Does NIST endorse you?
- Who sets the standard, who checks it, and who signs off?
- Who checks the checker, and what happens when a gate fails?
- Is the ruler circular?
- How do you stop a model from gaming or memorizing the questions?
- Did my model or agent get better or worse?
- What is calibrated today, and what is not?
- What is the IP and legal position?
Enterprise AI/ML leaders
Lean Six Sigma and I-O practitioners and consultants
- My client does not care about the science. What can I sell?
- How does this map to DMAIC and the capability indices I report?
- What counts as an item?
- What does level 9 against level 10 mean?
- What does it cost?
- How does output reach my client?
- Can I use this for selection decisions?
- Can we measure work that already exists?
Researchers
- Can the number be trusted?
- Is the ruler circular?
- Who checks the checker, and what happens when a gate fails?
- Who sets the standard, and where is the method published?
- Can I get the engine and data to replicate?
- What does level 9 against level 10 mean?
- How do you correct for judge severity on open answers?
- What is calibrated today, and what is not?
Answers
Can the number be trusted?
Every number carries its own stated standard error, so you see how sure we are, not only the score. A pre-registered precision gate for one judge's harshness passed on 2026-08-22; that judge has since retired, and the record is dated and public. That gate checks the method against itself. A check against expert judgment outside the method is the next exhibit, and it is not done yet. Uncertainty is reported in the style of the GUM conventions.
Was it the model, the judge, the data, the task mix, or my own test prompt?
We hold the other parts still so one can move. The question set is locked, the task mix is recorded and the grading prompt is held constant. Each judge's severity is estimated at each level and removed; that correction is modeled today, not yet gate passed. Your own test prompt is held constant, not measured, so we can rule it out but not size it. See how it works.
Did my model or agent get better or worse?
Re-measure on the same locked question set. A fingerprint check runs before every run, so any question change is flagged. Each judge's severity is measured at each level and removed; that correction is modeled today, not yet gate passed. A change smaller than the stated standard error is not a detectable change. Judges drift too: your judge changed; did your scores notice? Outside evidence of drift: Chen, Zaharia and Zou (2024).
How do you stop a model from gaming or memorizing the questions?
The question banks and calibration data stay closed, so public sources do not expose them to training. Each question set is frozen under a fingerprint, so any change is caught before a run. The model that writes a question never grades it, and no judge grades its own company's questions. We explain the design in contamination-proof evaluation. We do not yet claim a measured contamination rate.
Is the ruler circular, graded by the same kind of models it measures?
Partly a fair worry, and we answer it on the Evidence page. The levels are defined before any answer is graded, from a published developmental theory (Commons, 2008). Graders come from other model families than the question writers, and each judge's severity is measured and removed. What breaks the circle fully is a check against human expert judgment. That check is next and is not done yet.
My outputs are long or subjective, and there is no single right answer. How do you judge them?
We do not look for one right answer. We rate the complexity of reasoning each answer shows, on levels defined before grading, with partial credit (Masters, 1982). Each judge's severity is measured at each level and removed; that correction is modeled today, not yet gate passed. In the demo, judges from different model families take turns; the current roster is on the Judges page. Our published persuasion study: Barney, Wind and Krishna (2026).
What does a run cost?
Cost follows two numbers: how many questions a performer needs to reach the precision you set, and how many judge calls score them. Adaptive selection asks fewer questions when answers are informative. The cheapest capable judge goes first; a stronger judge on stand-by is in build. We aim to quote a run against your target precision before spend. For consultants, the unit is per engagement; for enterprise teams, per program.
How is this different from Trismik or the eval tools in my platform?
Trismik offers item response theory and adaptive testing for LLM evaluation, a serious approach we respect. Platform eval tools help you author tests and report pass rates. AIM puts people and AI systems on one ruler of reasoning complexity, estimates each judge's severity, and states a standard error on every number. Our method is published and our gates are pre-registered. Choose the tool that fits the question you need answered. Source: Trismik's description of its adaptive testing, Trismik blog, 2025-09-15, checked 2026-10-08.
What is calibrated today, and what is not?
The Hallucination and Persuasion scales both read "In calibration" as of 2026-10-08. Judge credentialing for levels 7 to 13 is in progress. Published judge ceilings, with pass counts and dates, are on the Judges page. A single run reports precision only; capability language needs repeated occasions. When a status changes, the scale page changes first.
How does output reach me? Do I log in, get a report, or call an API?
Today: a written report on request, and a recorded session you can replay end to end. The engine can also run inside your own environment, on request. A headless API and an MCP endpoint are planned and are not callable today (docs). We will not call anything live that is not.
Who checks the checker, and what happens when a gate fails?
We file each gate before the data arrives, on the Open Science Framework. When a gate fails, we report the failure as a finding. One already has: a 200-question wave did not pass the test for one shared dimension (filed 2026-07-20), while its other gates passed. We do not yet claim this catches a real production failure. A gate that can fail is what you are buying.
Can I bring my own judge and keep data inside my boundary?
The engine can run inside your own environment, on request, so less of your data needs to leave it. Bringing your own judge model is designed, not built: the plan is that you plug in your endpoint and we measure its severity before its scores count. No self-serve version is live today. Ask us about running it in your environment now.
Does NIST endorse you?
No. NIST does not endorse vendors, and we claim no endorsement, validation or compliance. AIM implements at measurement grade the latent-trait approach that NIST AI 800-3 highlights as a promising foundation for AI evaluation statistics. The report is public: NIST AI 800-3. We explain the link on our NIST page.
Who sets the standard, who quality-checks it, and who signs off?
The method is published as Chapter 3 of a De Gruyter volume (2024). We pre-register every gate and report failed checks as findings. We name no accreditation we do not hold. Our vocabulary follows the VIM, and our validity argument is organized around the AERA, APA and NCME Standards. Each claim on our pages carries its status: internally certified, model output, fact or proposed.
What does it cost us to keep deciding without this?
Model swaps, vendor choices and task hand-offs often rest on scores with no stated error. When a score moves, you cannot tell a real change from noise, a judge change or a prompt change. So you pay twice: once for the wrong call, and again to find out it was wrong. AIM puts a standard error on each number, so you know which decisions the evidence can carry. See the evidence.
My client does not care about the measurement science. Can you give me something I can act on or sell?
Sell the decision, not the method. Your client wants to know what limits the work, and which tasks go to a person, an AI or both. The output names the constraint, the measured options and a recommended pathway per task, with standard errors on the measures. No client has received one yet. Read the reasoning in our white papers.
Has anyone actually deployed this commercially? Can you show me a reference?
No. We have no named commercial deployment, and we will not invent one. What you can check today: published judge ceilings with pass counts and dates (Judges page), a recorded replay on a scripted test performer with 139 locked questions, and the published method. Start with the recording.
Our people already produce the work. Can we measure what exists instead of running a test day?
In design, yes. Work your people and systems already produce can be scored, so nobody stops for a test. Each piece of work enters one model and returns one estimate with one standard error. Every recorded run so far scores text; we have not run audio, image or video. The method has not yet been run on candidates or against job outcomes. Using it for decisions about people brings the same legal duties as selection.
How long will the security review take, and what will you need from us?
From us: the published method, the engine source, on request under a reviewer agreement, and a plain inventory of every data input, where it lives and which providers process it. From you: your questionnaire and your data rules. We answer item by item, in writing. We will not quote a timeline we cannot keep.
If we collaborate, what is the IP and legal position?
The engine is planned for Apache-2.0 release; reviewers get it on request under a written agreement; no public repository yet. Question banks and calibration data stay closed. What crosses between us goes in writing before any measurement runs, so you read it before you commit. We are not your lawyers, and this page is not legal advice. See our terms.
How does this map to DMAIC and the capability indices I already report?
AIM serves the Measure and Control phases. Each measure has a lower and upper spec limit and a target, and every performer location carries its standard error, so CpK-style capability can be read against your spec bands. Capability language needs repeated occasions: a single run reports precision only. Redesign work then follows DMAIC or DMADV. See transformation.
What counts as an item when you measure this way?
An item is one scored task with a stated difficulty on the ruler. Items are frozen under a fingerprint, so your score traces to the exact items, their keys and the grading rule. Where a role map exists, items come from tasks rated important in a job analysis. Our design for that step has not been run on human data yet. See the glossary.
Where does my situation sit on your scale, and what does level 9 against level 10 mean?
Each level is a step up in the complexity of reasoning the work demands, defined before any answer is graded (Commons, 2008). Level 10 work coordinates several level 9 lines of reasoning into one system. The published ruler spans levels 7 to 14; the 2026-08-22 precision gate covered levels 8 to 12. Every location carries its own stated error. No worked case is published yet. See the ruler.
Can I use this for selection or hiring decisions?
Not today. The method has not been run on candidates or validated against job outcomes, so it cannot carry a selection decision today. Selection use falls under the Uniform Guidelines and laws such as NYC Local Law 144, which add validity evidence and bias audit duties. We will tell you plainly when that evidence exists. Our validity argument is organized around the AERA, APA and NCME Standards.
Can I get the engine and data to replicate your work?
The engine is planned for Apache-2.0 release; reviewers get it on request under a written agreement; no public repository yet. Our gates and analysis plans are on the Open Science Framework. Production question banks stay closed. Published studies carry their methods: see Barney, Wind and Krishna (2026) and the method chapter. Ask us about reviewer access.