We measured our own judge at eight levels of task complexity. Two cleared. Here is the rule that follows, free to copy.
2026-09-22 · capability routing
Is Jev, a new AI model in beta, a good judge? It makes decisions instead of text: given work, it returns a verdict.
Spotty. Good at some levels of decisions, not at others.
We measured it at eight levels of task complexity. Two levels cleared the declared bar. Six did not. Where it clears, it is eligible for the work. Where it does not, it is excluded for that item, and you route around it. The rest of this post earns that answer: the mechanism, the numbers, and the rule you can copy.
If you stop reading here, hold two things. A judge is admissible only inside its measured range. And the measurement is a routing verdict, not a quality verdict.
Somewhere in your measurement and evaluation pipeline, a model is scoring work right now. Somebody picked it. If the pick was made by habit, by relationship, or by rotation, the number it produces has a defense problem: the honest answer to how that judge was chosen describes a procedure that never asked whether the judge could do the work. The number inherits the answer.
The scoring is the cheap part. The cost sits in the decisions made on the number: the release that shipped, the vendor that was renewed, the model that was promoted, each resting on a judgement from a judge whose behavior at that level was never measured. You do not need a new model to fix this. You need a decision rule, and this post gives it away for free.
Most teams that score with models run a round robin: the judges take turns. It is fair, it is cheap, and it spreads your exposure across providers, so no single vendor silently shapes every verdict. As a load balancer, it does its job, and nothing here is an argument against running one.
But a turn is not a measurement. The round robin decides who judges this item. It decides nothing about whether that judge can judge this item. Those are different questions, and the second one is the question your audit asks.
A judge is calibrated as a rater facet in one joint fit over a bank that carries the origin, so its severity sits on the same scale as the graders it will sit beside. That is the whole mechanism: the judge is measured like any other rater, on the same ruler, in the same coordinate, and its severity is disclosed with its standard error.
Because the bank carries the origin, that severity is not a free-floating preference. It is a position in the same coordinate the human graders occupy, so a contrast between the judge and a grader carries no dependency on where the scale was zeroed.
Severity is not what makes the range, though. Fit is what admits. A level belongs inside the range only where the fit holds under a frozen rule and a precision condition, and the range is the maximal run of admitted levels, re-derived by rule from any new fit rather than edited by hand.
Here is a test you can run on any number anybody shows you: change the offered menu of options and watch the number. A value that moves with the menu was never a location on a scale, and it will move again the next time the form of the question changes. That test costs nothing, and it separates a measurement from a reaction.
Now the numbers, which we would once have kept off the page. The levels are positions on the Model of Hierarchical Complexity (MHC) ruler, and the severity values on them are logits on the shared ruler. The calibration vintage is 2026-09-20. The declared fit band is 0.50 to 1.50, and it was declared before the fit was computed. The judge is Jev, and it is in beta.
Eight levels were measured. Two are admitted: level 8 and level 10. Their fit statistics sit at 0.88 and 0.58, inside the band, and their severity values sit at -0.09 and +0.10, close to neutral. The standard error at each admitted level is 0.04.
Six levels fail the band: 5, 6, 7, 9, 11, and 12. The failure is not marginal, and it is not a sample size problem. At level 7 the fit statistic reads 0.01, which means the judge echoes the reference so closely that it adds almost no independent information of its own. At level 11 the fit statistic reads 3.45, which means its departures from the reference are mostly noise. The remaining excluded levels fail from the low side of the band, one statistic or both sitting under it: they track the reference too closely to add much independent information. Two of the six, levels 5 and 12, also miss the precision target, with standard errors of 0.13 and 0.09 against a target of 0.07.
Severity runs from +0.43 at level 5, the harsh end, to -0.74 at level 12, the lenient end (above 0 = harsher), and the admitted levels sit near zero. A severity value is a disclosure about bias, not a grade for quality: it tells you which way the judge leans at a level where its fit already holds.
Read the fit numbers as a routing verdict, not a quality verdict. Nothing here says Jev is a bad judge. It says where Jev is usable, and the usable set is two levels wide; outside it the judge is excluded for the item, never demoted.
Every judge arrives with a different profile of the levels it can be trusted at, and the next judge will not arrive with the profile this one has. So the admissible range cannot be a setting somebody remembers to configure: it has to arrive with the credential, be read at selection time by the thing doing the selecting, and be re-derived whenever the underlying fit moves. A range held in a person's head is wrong the first week a new model ships, and models ship constantly. The range is part of the credential or it is nothing, and the selector consults it on every item rather than once at setup.
Four sentences. Copy them, use them, hand them to your team.
1. Use a judge only where its own measured range covers the interval you are judging.
2. Of the judges that qualify, take the cheapest and the fastest first.
3. Rotate the qualifying remainder rather than repeating the first.
4. A judge that does not qualify is excluded for that item, never demoted.
No paywall, no email gate, no contact form. The rule is the map, and the map is yours.
Copying the rule accomplishes nothing until every judge in your pool carries a measured range. The measurement is the work: each judge calibrated as a rater facet over the same bank, each range derived by rule from that judge's own fit. That is the part that cannot be copied from a post, because it is a measurement of the judges, and a measurement has to be run.
The runtime that consumes it is a selector over a pre-computed capability surface: computed offline across the item-by-judge space, versioned to the bank and the registry it came from, consumed as a lookup at selection time, and marked stale, failing loudly, when a recalibration touches a cell it covers. The surface has to be built before the selector. A selector with no surface has nothing to read.
Stated plainly: the measurement is certified today. Routing on it is the build, not a shipped claim. What is being built carries three commitments. Reliability is tracked as a posterior that updates with evidence, not a static snapshot. Availability and cycle time are read live, so a slow or offline judge is demoted in real time. And no judge is ever seated on an interval its own measured range does not cover.
These are early results. Every figure in this program is tied to a calibration vintage, and the vintage for every number above is 2026-09-20. Jev is a vendor-served alias in beta, so a provider release can change its behavior, entries go stale after 90 days or a version bump, whichever comes first, and stale entries are marked rather than deleted. A recalibration is expected, not exceptional.
Judgements from Jev have cost nothing so far, and the reason is that the vendor's listing for it is zero, captured 2026-09-20. A listing can move, and so can that fact.
No ceiling and no floor have been measured for any judge in this program, and the measured range is a routing rule, not a capability boundary. One judge is credentialed today. The rest of the roster is not yet, and the process that made the one turnkey is the point.
We measured our own judge at eight levels, and the rule admits two of them. The other six are published in the figure above, failure values beside the successes. More rows cannot clear them: what fails is the fit, not the sample size, and the fix is new items, not more data.
We publish the failures because a house that shows its own numbers, both columns, is demonstrating the discipline it sells. The claim is not that capability-conditional judging was invented here. The claim is that we measured the edge and we can bound it, which a rotation cannot do and a static table cannot do, because neither re-derives a range from a fit.
Scoring starts free and needs no card. Price is set by the standard error your decision requires, and every engagement is scoped and quoted before it is priced. The pricing surface is at https://xlnc.co/pricing.
Every figure in this post is tied to a calibration vintage. The API and MCP server do not exist yet, and no promise of them is made here.
You will pick a judge for the next item either way. The rule costs nothing, and it turns that pick into a measurement. Running it at scale, on every item, against measured ranges that arrive with each credential, is the product.
Calibrated, traceable measurement for decisions that have to survive scrutiny.
Founding cohort: priority access, and a seat to shape the instrument. Email is enough.