Your judge changed. Did your scores notice?
The model that scores your work can change under the same name. Here is how to tell a worse product from a changed judge, and the checks we are piloting.
The pass rate moved and nobody touched anything
When a pass rate moves and nobody changed the product, check the judge: providers can change the model behind a name, and Chen, Zaharia and Zou (2024) found GPT-4's accuracy at telling prime from composite numbers fell from 84% to 51% between March and June 2023 with the name unchanged. The check we are piloting re-measures the judge on locked questions and dates every reading.
Picture a Tuesday. Your team shipped nothing on Monday. The prompts are the same, the dataset is the same, the rubric is the same. Your evaluation dashboard says the pass rate on last week's release rose four points overnight.
Did the product get better? Nobody changed the product. So what changed?
The honest answer is often the one nobody checks: the judge. The model that scores your outputs is served by a provider, and the provider can change what sits behind the name. Your dashboard has no field for that, so the change arrives looking like progress.
It happens to careful teams
This is a documented pattern, and the teams it has hit are among the most careful in the field.
The clearest early record comes from Stanford and Berkeley. Chen et al. (2024) asked the same questions of the same named services in March 2023 and again in June 2023. GPT-4's accuracy at telling prime from composite numbers fell from 84% to 51%. Over the same three months GPT-3.5 moved the other way on the same questions, from 50% to 76%. The model names never changed.

Figure. Four tasks from Chen et al. (2024), each scored on the same named service in March and June 2023, on a full 0 to 100% axis. Every value is printed in the text or figure captions of the paper's preprint (arXiv:2307.09009v3, Sections 3.1, 3.2, 3.3 and 3.5). On three of the four tasks the two models moved in opposite directions. A judge built on either one would have changed its verdicts without changing its name.
The drift was not only in accuracy. On the prime question, GPT-4's average answer shrank from 638.3 characters in March to 3.9 in June (Chen et al., 2024, Section 3.1). Anyone scoring GPT-4's reasoning in March was, by June, scoring a bare verdict.
Serving changes can move quality without any new model at all. Anthropic's postmortem describes three infrastructure bugs that degraded Claude's responses between August and early September 2025 (McAllister, 2025). One routing bug touched 0.8% of Sonnet 4 requests when it began on 5 August and 16% in the worst hour, on 31 August. Users noticed first. Anthropic's own account of why it took so long is plain: "The evaluations we ran simply didn't capture the degradation users were reporting."
Behavior can move in a direction that flatters the reader. OpenAI rolled out a GPT-4o update on 25 April 2025 and rolled it back days later because the model had become sycophantic, and its follow-up, as quoted by Willison (2025), says its offline evaluations and A/B tests had looked good.
Three providers, three kinds of change: a model update, a serving fault, a behavior shift. In each case the provider's own checks were built by people who knew the model better than any customer does. If their evaluations could miss it, what is the chance that a judge in your pipeline is being watched more closely?
Same outputs, opposite verdicts
A pass rate reports what the scorer checked, and the scorer decides the story.
Chen et al. (2024, Section 3.5) show this on code. Scored as "runs as delivered," GPT-4's June code collapsed: directly executable answers fell from 52% to 10%. Scored as "passes the tests once the extra formatting is stripped," the same June code improved, from 52% to 70%. One set of outputs, two scorers, a 42-point fall or an 18-point gain. A dashboard shows one of those numbers and hides the other.
Movement alone is not drift, either. The same paper ran GPT-3.5's March version twice on the same opinion survey: 2.8% of its answers disagreed with themselves. Between March and June, 27% of its opinions changed (Section 3.4). The first number is the instrument's wobble. The second is a change. A monitor is only useful if it can tell them apart, and it can only do that if it knows the size of the wobble.
A judge that drifts is a ruler that stretches
A score from a model judge is a reading taken with an instrument. If the instrument changes, the reading changes, and the thing being measured may not have moved at all.
Li (2026) states the problem in one line: "every drift alarm is ambiguous between a worse product and a changed judge." A tape measure that stretched an inch overnight makes every board look shorter. You can re-cut the boards, or you can check the tape.
The cost of not checking lands on decisions, not on dashboards. A lenient shift passes work that should have failed, and a severe shift blocks a release that was fine. The release ships or stalls, the vendor is renewed or dropped, the model is promoted or retired, and each call rests on a number whose instrument quietly moved. When somebody later asks why the number changed, "the provider changed the judge and we did not notice" is a hard sentence to say to an auditor, and a harder one to say to the person whose work was scored.
Why the usual checks miss it
Most teams already run one of four checks. Each answers a narrower question than the one a changed judge raises, and the published numbers show where each one runs out.
- Manual spot checks. At 0.8% of requests, the rate at which Anthropic's routing bug began (McAllister, 2025), a reviewer reading 100 responses would expect to find fewer than one bad one. The fault grew to 16% in its worst hour, and user reports, not evaluations, raised the alarm.
- Vendor-side evaluations. Anthropic's evaluations "didn't capture the degradation users were reporting" (McAllister, 2025). OpenAI's offline evaluations and A/B tests looked good before a rollback (Willison, 2025). These checks protect the vendor's product in general, and your rubric is not in them.
- Raw pass rates. The same June code scored 10% or 70% depending on what was checked (Chen et al., 2024). A pass rate with no record of its scorer cannot say which kind of change you are looking at.
- A single judge with an alarm and no baseline error. In a preprint, Li (2026) reports that a common rolling z-test raised false alarms on 75% of streams that had no drift at all. An anchor set that the judge re-scored on a schedule caught a silent version bump in 60 of 60 runs, with no judge change blamed on the system.
What the last result shows is the common thread. A drift check needs a fixed reference the judge is measured against, and an error bar on that measurement. Without both, no alarm can say whether it saw a real change or a wobble.
Change is scheduled, too
Silent change is the hard case. Announced change is constant, and it breaks things just as surely.
OpenAI's deprecation policy promises at least six months' notice for generally available models and as little as two weeks for previews (OpenAI, n.d.). Anthropic's page lists eight retirement announcements between January 2025 and June 2026, with at least 60 days' notice for publicly released models (Anthropic, n.d.).
Settings move as well as models. The same Anthropic page records that, on Claude Opus 4.7 and later, setting temperature to a non-default value returns an error. A judge that was measured at one temperature cannot be run at that temperature on the newer model, so "the same judge" is no longer the same configuration.
It happened to us too. On 20 September 2026 the maker of one of the models we score with released a new version. In our pilot probe of 19 August 2026, that judge read 1.75 stages more lenient than the reference (95% CI 1.63 to 1.86, n = 1,336). Whether the new version leans more, less or the same, nobody outside its maker knows, and until it is re-measured, neither do we. Our record for that judge was marked stale the same day, and the judge is not used until it is measured again.
What drift monitoring checks
A drift check compares a judge today with the same judge at its baseline, on the same ruler, and asks whether the difference is larger than the measurement error. A drift check with no standard error is a guess with a chart attached.
These are the tests we are piloting. Each fires only when a change is both statistically and practically large, and the tests for one judge are corrected together so that running many of them does not manufacture alarms.
- Severity shift. The judge's severity is re-estimated and compared with its baseline, and the shift is divided by the combined standard error of the two readings. It fires when that ratio passes 2.58 and the shift is at least 0.10 on the scale.
- Fit degradation. The judge's fit statistics are re-computed. It fires when the fit leaves the band declared before the data were fitted, or worsens by more than a set amount.
- Reliability. The judge re-rates items it has rated before. It fires when agreement with its own earlier ratings falls below the floor.
- Availability and latency. Success rate and response time are tracked as series. It fires on a real drop in availability or a real slowdown, with latency counted per unit of output so that a wordier model is not mistaken for a slower one.
- Served-model change. Any change in the model identifier the provider reports, or in the settings the judge runs with, fires every time, with no statistics needed.
To see how the severity rule reads, take an illustrative case in which both readings carry a standard error of 0.04 logits. The smallest shift that fires is then 0.146 logits, because 2.58 times the combined error of 0.057 exceeds the 0.10 floor. A shift of 0.08 stays quiet however often it recurs, and a shift of 0.26 fires with its size and error attached. A monitor that alarmed on every wiggle would be switched off within a month. One that alarms with a number and an error bar gives a decision something to stand on.
The last test depends on a settings fingerprint: a hash of everything that can change a judge's response, from the exact model identifier and temperature to the system prompt, the rubric version and the scoring code. A measured judge is a model plus that exact configuration. Change one setting and you have a new, unmeasured configuration, and the fingerprint makes that visible before a single score is written.
Three checks you can start this week
None of these needs us, and each one narrows the gap between a changed judge and a changed product.
- Log the served model identifier and every judge setting next to every score. A score whose judge configuration you cannot name cannot be compared with last month's.
- Keep a small fixed set of already-scored items and have the judge re-score them on a schedule. If those scores move, the judge moved, because those items did not.
- Run the judge twice on the same items and record how often it disagrees with itself. That is your wobble. Chen et al. (2024) found 2.8% for one model; yours will differ, and any alarm threshold below it will fire on noise.
Updates, and where drift monitoring sits in it
Drift monitoring is one line in Updates, a product we are developing and piloting. As designed, Updates would deliver three things, none of them a finished service today: improved parameter estimates for each judge as new data arrive, the reliability, availability and latency readings the judge-selection rule uses to seat a judge on an item, and the drift checks above. In the design, a judge that drifts is set aside for the work it no longer covers until it is re-measured, and it is excluded rather than demoted.
Updates is in development and is early evidence, not yet certified. Per-level judge credentials are the formal record of the levels at which each judge may be used. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way. Every figure we publish carries its calibration date.
Measure the instrument, not only the work
Every team that scores with a model is trusting an instrument it did not build and cannot inspect, served by someone who can change it. That trust can be checked. The judge can be measured, the measurement can be repeated, and the difference can be read against its error. We think every score that decides a release, a vendor or a model deserves that, and we are building toward the day it is routine: a field where a judge's reading carries its date and its error bar as plainly as a lab result does.
The Design-Partner Readout is the place to start. Send a redacted slice of the evaluation data you already keep, traces included, and in about 30 days you get a written early read of how much of a score change came from the model, the judge, the dataset, the task, and the prompt. If your pass rate has ever moved when nothing on your side did, that is the question it answers. Ask for one at xlnc.co/solutions.
Different teams need different things, and each of these is offered as early access: the Design-Partner Readout for one score change you need explained, a hosted plan designed to score through us, a self-hosted harness designed to keep your data on your side, a custom pilot built around one decision, or Updates, which we are piloting to keep judges measured as providers change them. Tell us which fits yours at xlnc.co/solutions.
This is early evidence: every reading in this program is tied to a calibration date, and the monitoring described here is being piloted, not sold as a finished service.
References
- Anthropic. (n.d.). Model deprecations. Claude Platform Docs. Retrieved September 28, 2026, from https://platform.claude.com/docs/en/about-claude/model-deprecations
- Chen, L., Zaharia, M., & Zou, J. (2024). How is ChatGPT's behavior changing over time? Harvard Data Science Review, 6(2). https://doi.org/10.1162/99608f92.5317da47
- Li, Y. (2026). Who drifted: The system or the judge? Anytime-valid attribution in LLM evaluation pipelines (arXiv:2606.15474) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.15474
- McAllister, S. (2025, September 17). A postmortem of three recent issues. Anthropic. https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues
- OpenAI. (n.d.). Deprecations. OpenAI API documentation. Retrieved September 28, 2026, from https://developers.openai.com/api/docs/deprecations
- Willison, S. (2025, May 2). Expanding on what we missed with sycophancy. Simon Willison's Weblog. https://simonwillison.net/2025/May/2/what-we-missed-with-sycophancy/