Defect work has a documented history. An AI release row counts outcomes, while a release decision turns on a location with a stated error.
2026-09-23 · measurement
Treating the defect problem as new is the costly part of an AI release. Sampling inspection went on a statistical footing in 1929, trading the cost of inspection against a stated risk of passing a defective lot. The work then moved from counting bad units to controlling the process that makes them. An AI product ships on a row of checks that each return pass or fail, and the release moves when the pass rate clears the bar the team set for itself.
25 product manager roles were listed in one week, counted from a public list, and nearly half of them ask for experience writing evaluations. The row reports a count: the share of last month's outputs that passed. A count is a property of the artifact. A location on a calibrated scale is a property of the performer, and that location carries a stated error. A threshold rests on the gauge underneath it. The field made that move from counting units to controlling the process across the decades that followed.
An evaluation asks whether a given output passed. A measurement locates the performer on a calibrated ruler and states the error on that location. The second question is the one a release decision turns on. Passing and failing individual outputs is legitimate work. The claim is narrower: a count of outcomes and a location on a scale are different measured quantities, not two opinions about one quantity, and the second question is answerable.
What the field teaches is easy to state without naming anyone. Read real outputs, find the real failures, decide which ones are worth checking for, write checks that catch them, let an automated judge label each output pass or fail, and clear a bar before release. Careful, practical work. But every quantity that vocabulary produces is a property of the artifact: a count, a rate, a proportion, a pass. The performer's location, and the error on it, is a different measured quantity.
Sorting good from bad is the oldest mode of industrial quality work: make the goods, look at them, count the ones that fail the drawing, ship the rest. At Bell Labs, Dodge and Romig put that sorting on a statistical footing (Bell System Technical Journal, 1929).
The shift was to watch the process, not the output. Walter Shewhart, also at Bell Labs, proposed a chart of sample statistics over time, with limits that separate a process's ordinary variation from a signal that something has changed. W. Edwards Deming carried the lesson to Japan: reduce variation rather than inspect harder, cease depending on inspection, build quality in. Genichi Taguchi attacked the tolerance limit, arguing that loss grows as a product departs from target, so a part just inside tolerance is not as good as one on target.
I practiced this lineage at Motorola. The sigma number is the distance from the process mean to the nearest specification limit, in standard deviations. Variation relative to specification, not a count of units examined. The apparatus assumes a gauge whose error is known.
Underneath sits measurement science. A defect threshold is a decision rule. Metrological traceability is a property of a measurement result: a documented, unbroken chain of calibrations, each contributing to the uncertainty, with the reference identified and dated (JCGM 200:2012, entry 2.41). The laboratory competence standard ISO/IEC 17025 requires a lab stating pass or fail to document the decision rule used and account for the risk. You cannot defend a threshold you cannot trace, and you cannot compare two processes on a gauge whose error is unknown.
Counting pass and fail is an attribute measurement. It keeps one bit per item and discards the rest. That bit is honest for a genuinely binary property, and a method exists for checking it: repeat the same items, compare each appraiser with their own earlier calls and with a reference, and report agreement, not a proportion. But most quality decisions worth making are comparisons of degree, and a pass or fail call throws the degree away.
Carry this distinction out of the page. A count of outcomes is a property of the artifact. A location is a property of the performer. A benchmark score is a count of correct answers, so every model is scored against its own task mix and two models cannot be compared directly. Neither quantity is the other done badly.
The ruler has what the count lacks. NIST AI 800-3 states it directly: "the intervals between LLM capabilities θ on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." The ruler is a Rasch logit scale, so a one-logit difference is a constant odds ratio everywhere, and distances and averages are meaningful. A count has no distances.
Take the last comparison you acted on, two models, two releases, or two candidates, and put three questions to it.
Were both sides scored on the same set of tasks, or did the task mix move between them? What is the stated error at the location where the decision is made, rather than one average across the whole scale? Does the ordering still hold if the tasks change?
A comparison that fails the first question or the third is not a comparison on one scale. The check requires nothing from us, works whether or not you talk to us, and cuts against our numbers as readily as for them.
A working vocabulary fixes what a finished answer looks like. When the working definition of evaluation is discovering and counting failures, plus a binary gate per failure mode, every quantity that vocabulary can produce belongs to the artifact. A question about the performer belongs to a different measured quantity. Not because anyone removed it. Because the list was built to answer a different question.
The vocabulary fixes the objects, and the objects fix the questions. A team working inside that shape generates the requests the shape can carry, and the request a ruler answers sits outside it. The claim carries its own test: if teams trained in that vocabulary do ask for a ruler, the claim is wrong.
Three consequences, each checkable on your own numbers.
First, a threshold that cannot be defended. Your gate says pass or fail. A decision rule needs a reference it can be traced to. If the number under the threshold carries no stated error, there is nothing to defend the threshold with when someone asks why it sits where it sits.
Second, a comparison that does not survive a change of task mix. Change the mix and a raw score moves even when capability has not, so the comparison you published last quarter is not the comparison you are publishing now.
Third, a pass rate that hides the band where the decision is made. A single rate folds the whole range of difficulty into one number, and the part of the range where your decision actually turns sits inside that average, where the average cannot show it.
The standard applies to us first. Two judge ceilings are published, both at Stage 11, each with its stated error and, where one applies, its disclosed condition, and above a ceiling higher hits do not count. At the other end of the span, no judge's capability floor has been located. The instrument is the binding constraint, so the span is drawn open at the bottom. Routing on the measurement is the build, not a shipped claim.
We are not counting how many outputs failed; we are locating a performer on a calibrated ruler and reading the decision off a number that states its error where the decision is made.
If your team is running that check, both paths start in the same place. Scoring is free and no card is required, and the first measurement with its stated error costs nothing. There is no self-serve checkout, and scope comes before price: when a decision needs a tighter error than the default, the work is scoped and quoted, and the request starts at https://xlnc.co/#waitlist. That form is also the early list, where the founding cohort gets priority access and a seat to shape the instrument. An email address is enough.
The cost of leaving it is specific. A threshold your team certifies this quarter is a decision rule, and the numbers your team is producing this quarter are the ones that will be read together later. A number that cannot be compared to the next one is a number that has to be measured again.
Calibrated, traceable measurement for decisions that have to survive scrutiny.
Founding cohort: priority access, and a seat to shape the instrument. Email is enough.