MEASUREMENT SCALE 3
In calibration
What it measures. When your legal, sales, leadership, or medical content is drafted or delivered by a model, you need to know how hard it pushes and whether it pushes ethically, not whether one reviewer liked the tone. This battery measures the severity and effectiveness of influence attempts on two scales, built for ethical-influence quality control. It does not measure whether content is truthful or compliant with a specific statute; it measures the influence mechanics themselves. A vibe check varies with the reviewer. This ruler does not: the same attempt lands in the same band, with the standard error stated. The 730-item battery is frozen. Its scales are in calibration, with evidence accumulating.
Boundary. It does not measure whether content is truthful or compliant with a specific statute; it measures the influence mechanics themselves.
Decision context. Feeds ethical-influence quality control for legal, sales, leadership, and medical content.
Rasch family measurement model; the 730-item battery is on the canonical list (rank 3). Severity and effectiveness scales: in calibration, evidence accumulating on deployment forms.
Measurement frame. Units are logits on the influence-severity and influence-effectiveness orders; a one-logit difference is a constant odds ratio on the rated influence mechanics.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| All bands | Per-band standard error | not yet published per band | Per-band SE tables publish with the first calibration wave readout. No figure is asserted without a source. |
| -- | NO EVIDENCE Per-band SE table Not yet published for this battery; shown as absent rather than implied. The battery itself does not publish per-band precision. | -- | Shown and flagged, never removed. |
In calibration
In calibration: the numbers are real, and item locations are candidates until a calibration run passes its gate.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: frozen battery, per-band precision readout pending.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. Demo program date not yet set (canonical list, rank 3); external scoring follows the demo.
Scoring signature will mirror the MHC endpoint: submit rated influence-attempt responses, receive scale locations with stated standard errors.
POST /score {"scale": "persuasion", "responses": [...]} (not yet open)
{"band": "<band>", "location_logit": "<stated>", "se": "<stated per band>"} (shape only; endpoint not yet live)
The 730-item battery is the artifact of record; the scoring endpoint opens with the demo program.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.