XLNCXLNC Watch it measure

Blog

Claude can now build your evals. Who measures the judge?

Anthropic's new eval and hillclimbing guidance gets the judge's independence right. Here is what a calibrated judge adds, with early data.

The puzzle

On 28 September Anthropic published guidance on building evaluations with Claude Code and improving an application against them, and added two commands to go with it. One builds an eval from your codebase. The other hillclimbs on it, one change at a time.

Read the guidance closely and a puzzle appears. A team that makes tools for writing software produced a playbook whose working vocabulary is measurement science: an independent judge, a confidence interval on the score, a check that the grader gives the same verdict twice, a noise floor, a warning about headroom. Why did a coding tool end up there?

Our reading is simple. The moment an agent iterates against a score, the score stops being a report and becomes the target. Every flaw in the instrument gets optimized. Questions that used to sit in a statistics appendix now decide whether a change ships. And if the score is the target, the judge behind the score is the next thing to check.

That is good news. It means the field has arrived at the right questions. This post is about the next one.

What the guidance gets right

Credit first, and specifically.

The article says the judge "should not be the model you are testing." It reports the score with a confidence interval. It runs the grader twice on the same output to check that verdicts hold. Before round one of a hillclimb it compares the smallest improvement worth acting on with the noise. It holds out cases to catch overfitting, and it warns when a score passes about 95 percent, because little headroom is left. If a result is within noise, the workflow says so and recommends against merging.

Each of these protects you, and most eval setups in the wild have none of them. If you use the new commands, you are better protected than you were last month.

Where a playbook stops and a measurement system continues

Each point below extends the guidance. A playbook has to stay short; a measurement system does not.

Independence names the judge; it does not measure it. The article has Claude grade a handful of cases and ask whether you would score differently. That is a useful spot check. A calibration would report the judge's severity and fit at each level of difficulty in your task, each with a standard error.

A stable judge can be consistently wrong. Rerun agreement finds noise. It cannot find bias. In our pilot data, one judge fit the measurement model as expected at level 8 of the Model of Hierarchical Complexity (infit 0.97) and read 0.13 one level down at level 7. That is overfit: the judge was so predictable that it can separate whole stages but not finer steps, so we tag that level as coarse grain only. Our fit rule is one-sided. A level fails when it underfits, meaning erratic ratings past a set limit; overfit is a tag, not a failure. Fit is not reliability, and rerun agreement does not measure fit, so it could not have told those two readings apart.

A noise floor needs a source. Checking noise against the smallest meaningful improvement is exactly right. What we could not tell from the article is whether that floor includes error from the judge itself. In measurement terms, the smallest detectable change comes from the standard error, and for a judged eval that includes the judge's severity, which moves from level to level.

Percent correct is a count, not a ruler. On the article's own support-ticket benchmark, 74.4 percent and 88.9 percent belong to different model settings on the same 44 tickets. Equal gaps in percent correct do not mean equal gaps in capability, and "saturated" is a threshold on a percentage, not a location on a scale. Placing cases and models on one interval scale is what a Rasch model does, and it turns headroom into something you can read off.

Human reviewers have a severity too. Practitioners such as Hamel Husain rightly insist on looking at the data before writing evals. We agree. The measurement point is that a human reviewer is another rater with a harshness of their own, and it can be modeled and corrected as a judge's can.

Models are jagged. So are judges.

The article warns that model capability is jagged, and that hand-picking the cases today's model fails measures that model's failure fingerprint. We would extend the warning one step. The judge has a fingerprint too.

We measured five LLM judges at levels 5 through 12, 40 judge-by-level cells. Thirty-seven gave usable estimates. Twenty-one were admitted, and 17 of those only at coarse, whole-stage grain. No judge was admitted at level 5, 11 or 12. The most expensive judge, Kimi-k3 (Moonshot AI), was not admitted at a level where three cheaper judges were. Severity moved above and below zero from level to level, mostly within the error bars. The full write-up is in The same judge, one level apart.

These figures are early evidence: five judges, one ruler, calibrations dated 2026-09-20 and 2026-09-23, under a procedure we are now replacing with a stricter one. They show that a judge's quality varies with the level of difficulty. They do not show that any named model is a poor judge in general.

A white graphic titled 'Four questions to ask your eval judge. What a measurement system reports.' Four rows. One: a credentialed judge, with severity per level and 95 percent error bars for one judge at levels 5 to 8. Two: an error bar and a smallest detectable improvement, marked illustrative. Three: fit to the model by level, with a one-sided acceptance zone bounded by an underfit limit, 0.13 at level 7 tagged overfit (coarse grain only) and 0.97 at level 8 for one judge. Four: difficulty on a shared ruler, with no judge admitted at levels 5, 11 or 12. Early evidence, not yet certified. Not affiliated with or endorsed by Anthropic.

What it costs to skip this

A hillclimb loop that optimizes against an unmeasured judge can raise the judge's score while quality in production stays flat. The team ships a change it cannot defend, and pays again to rerun once the problem surfaces. A team should not merge a change it cannot defend.

The train and test split in the article protects you from overfitting the cases. It does not touch a judge that is lenient at the one level where your hardest cases sit. That exposure is our reasoning, not a measured result, and we would like to test it on real hillclimb runs.

Before your next round, ask whoever supplies your judge for three things: its fit at each level you use it, its severity at each level with a standard error, and the levels where it was measured and did not qualify.

Try it on your own judge

If you run LLM judges in your evals today, tell us which ones and at which levels of difficulty. Join the waitlist or book 20 minutes. Early access is by conversation. We will walk through your judge against the three items above, and we will say plainly what we could not measure.

What comes next: an AIM MCP server

Our goal: every judged score in your loop carries an error bar, and every judge reports where it qualifies.

To get there, an AIM MCP server is on our roadmap, and we plan for it to be integratable with Claude as an evaluation tool next quarter. A judge you already use, or one we credential, would report its profile and return each verdict with an error bar inside the loop you have already built. We also plan a plain API alongside it, for eval runners that execute in CI.

This is a plan. It depends on our own certification and security review, and on what the tools themselves allow. XLNC is independent of Anthropic, with no affiliation, endorsement or partnership. Anthropic, Claude and Claude Code are names of their owners, used here to identify the guidance we discuss.

Until then, the offer above stands.

This is early evidence: every AIM figure in this post comes from calibrations dated 2026-09-20 (Jev) and 2026-09-23 (the other four judges), and the credentials behind them are being redone under a stricter procedure. Descriptions of Anthropic's guidance are our summary of the article published 2026-09-28 and are not a statement of how any Anthropic product works internally.

Frequently asked questions

Does Anthropic's eval guidance already cover judge quality?

It covers judge independence, a confidence interval on the score, a rerun check on the grader and a noise check before a change is accepted. Those are real protections. What it does not describe is measuring the judge itself: its severity and fit at each level of difficulty, with a standard error.

Is a stable LLM judge a valid one?

Not necessarily. Stability means the same verdict on repeat. A judge can be stable and biased at one level of difficulty. In our pilot, one judge fit the model at one level and overfit one level below it (too predictable, usable at coarse grain only), and rerun agreement does not measure fit.

Will XLNC work with Claude?

An AIM MCP server is on our roadmap, planned to be integratable with Claude as an evaluation tool next quarter. It is a plan, and XLNC is not affiliated with or endorsed by Anthropic.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.