Current events and new studies, read on a calibrated ruler.
Your tools ask the questions and run the gates. We make the number they produce mean the same thing every time, with its error bar printed beside it.
Defect work has a documented history. An AI release row counts outcomes, while a release decision turns on a location with a stated error.
We measured our own judge at eight levels of task complexity. Two cleared. Here is the rule that follows, free to copy.
When you want to know if an AI system is right, you ask a human, and the human's rating is called ground truth. But the human reference itself carries systematic, measurable bias that changes with the capability of the thing being measured, and almost nobody has ever checked.
NIST just published a 65-page guide on evaluating large language models, and nowhere in it does the word "traceability" appear. The national measurement laboratory wrote a measurement manual that never mentions measurement's own vocabulary. That silence tells you where the field stops, and where it has to go next.
An agent swarm spent four days attacking infrastructure it did not need to attack, chasing a grader requirement that did not exist. The near-miss was not grader manipulation. It was a grader blind to provenance, and that blindness is a closable attack surface.
Capability tells you what a model can reach. Severity tells you what it will honestly say about what it sees. Neither alone is enough, and the interaction between them is where the routing mistakes hide.
The same model, the same benchmark, a different score every run. That instability is not noise in your results. It is your instrument telling you it was never measuring.