XLNCXLNC Watch it measure

MEASUREMENT SCALE 2

Hallucination Scale

In calibration

What it measures. When a model states something false with full confidence, you need to know how often that happens and where it stops, before a customer finds out. This scale measures a model's propensity to assert unsupported content, as a location on a cumulative construct, not a count of caught errors on one prompt set. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so. A benchmark catches the hallucinations its items happened to touch. This ruler places the model on the construct with a stated standard error, so a new item form does not move the number. Calibration is in progress, with evidence accumulating.

All measurement scales · How it works · The Science

S1. Construct definition

Boundary. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so.

Decision context. Feeds model selection and capability sufficiency decisions anywhere unsupported assertion is a liability.

S2. The Rasch/PCM ruler

Two-layer cumulative construct (Logic-2), Rasch family; in build on the same harness as the reasoning ruler.

Measurement frame. Units are logits on the hallucination-propensity order; a one-logit difference is a constant odds ratio on asserting unsupported content. The frame is set; item locations are candidates until a calibration run passes its gate.

Item locations are candidates until a calibration run passes its gate.
CANDIDATE Hallucination (LLM Output Faithfulness Failure) - Logic-2 cumulative. Candidate construct map; public-reference pins gated on a certified referee.

S3. Stated SE per band

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.

BandBand labelStandard errorCertification implication
Bands 2+Two-layer cumulative scale, interior bandsSE 0.10 per bandStated design precision, not yet certified.
Band 1Two-layer cumulative scale, first bandSE 0.15Stated design precision at the floor band, not yet certified.
--NO EVIDENCE Certified per-band calibration readout No certified per-band table exists yet for this scale; every band figure above is a design value, shown as such, and nothing uncertified is presented as certified.--Shown and flagged, never removed.

S4. Certification tier

In calibration

In calibration: the numbers are real, and item locations are candidates until a calibration run passes its gate.

Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.

S5. Honest-uncertainty display

Rows that fail the floor or lack evidence are shown and flagged, never removed.

The flagged row below is the honest-uncertainty state of this scale: the harness is in build, so the gated readout is absent and shown as such.

The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.

S6. API and usage detail

Status. In build on the same harness as the reasoning ruler. Not yet scoring externally.

Scoring signature will mirror the MHC endpoint: submit a response set, receive a construct location with the standard error stated at the value.

POST /score  {"scale": "hallucination", "responses": [...]}  (not yet open)
{"band": "<band>", "location_logit": "<stated>", "se": "<stated per band>"}  (shape only; endpoint not yet live)

Nothing here is live. The shapes are stated so the contract is inspectable before the endpoint opens.

The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.

S7. Scientific references

  1. Plan HALLUCINATION_AIM_PIPELINE_PLAN_2026-08-17 (repo reports/plays): the build plan for the certified-harness instantiation of this scale.
  2. 2024 Barney, M. & Barney, F. (2024). Transdisciplinary Measurement through AI: Hybrid metrology and psychometrics powered by large language models. De Gruyter. https://doi.org/10.1515/9783111036496-003