XLNCXLNC Watch it measure

Blog

Fresh instruments on demand: the path to contamination-proof evaluation

Benchmarks leak into training data, and a leaked benchmark measures memory. AIM's pipeline is built to generate fresh, calibrated items at measurement time. Here is what exists today, and what does not yet.

Picture a model that scores well on a public benchmark. Then someone writes a new set of questions, same kind, same difficulty, never published. Suppose the score drops. Nothing about the model changed between the two runs. The only difference is that the model had never seen the second set.

That gap has a name: benchmark contamination.

Short answer. Benchmark contamination happens when evaluation items appear in a model's training data, so the score measures memory instead of capability. Refreshing a benchmark by hand delays the problem and does not end it. The durable fix is to generate new items at measurement time and place each one on a calibrated scale before it counts. XLNC's AIM pipeline is built to do this along two pathways. The code exists. It is early evidence and not yet certified. We do not claim contamination-proof evaluation in production today.

Why a leaked benchmark stops measuring

A benchmark is an answer key. Once it is public, the clock starts. Web crawls pick it up, forum posts quote it, and a model trained on that crawl can meet the items before it meets the test. From then on, a high score has two possible explanations, and the score alone cannot tell you which one you are looking at.

Claim. Benchmark contamination is documented in the research literature, not just suspected.

Evidence (peer-reviewed).

  • Sainz, Campos, García-Ferrero, Etxaniz, López de Lacalle, and Agirre (2023), "NLP Evaluation in Trouble: On the Need to Measure LLM Data Contamination for Each Benchmark," Findings of EMNLP 2023 (aclanthology.org/2023.findings-emnlp.722). They show that contamination leads to overestimating a contaminated model's performance, and argue it must be measured benchmark by benchmark.
  • Golchin and Surdeanu (2024), "Time Travel in LLMs: Tracing Data Contamination in Large Language Models," ICLR 2024 (proceedings.iclr.cc). They give a method for detecting that benchmark items were seen in training, and report contamination in a widely used frontier model on several public datasets.

The usual response is to write a new benchmark. That works for a while. It also costs expert hours every cycle, and it creates a second problem: each hand-written set is a new instrument with its own unknown difficulty. Scores from the old set and the new set do not sit on one scale, so a trend line that crosses the refresh is comparing two different rulers.

We made the scale argument at length in Benchmaxxing is a measurement failure. The short version: raw scores depend on which items happened to be picked. Contamination makes that dependence worse, because the items that happened to be picked are also the ones most likely to have leaked.

What changes when the item did not exist yesterday

An item written at measurement time cannot be in anyone's training data. That is the whole appeal of fresh items, and it is half the answer.

The other half is calibration. A fresh item of unknown difficulty produces a score of unknown meaning. If today's items happen to be easier than last month's, the model looks better and nothing improved. So a fresh item has to land on the scale, with a stated error, before it is allowed to score anyone. Freshness without calibration trades a leaked ruler for an unmarked one.

Generating fresh test cases is not new, and it is useful work. Several evaluation tools already do it: DeepEval's Synthesizer builds synthetic test cases from documents or from scratch, and Patronus generates adversarial test suites. If you use them, keep using them. The part we have not seen described in public materials is the next step: placing each new item on a calibrated difficulty scale, with a standard error, before it is used.

AIM's design puts both requirements in one pipeline: generate the item, then locate it, then use it.

Two pathways, one locating step

Capabilities are not all built the same way, so AIM's pipeline has two seed pathways. Both finish in the same locating step.

Hierarchical constructs

Some capabilities are ordered by complexity: each level coordinates the actions of the level below, as in the Model of Hierarchical Complexity (MHC). For these, a Bayesian seed uses anchor items from a construct map as the frame and estimates where a new item is likely to sit.

Cumulative constructs

Other capabilities are built step on step: each level adds to what the previous level already required. For these, a sequential seed (S-SBS) gives the new item a starting position within its level.

The locating step: inverted CAT

Ordinary computerized adaptive testing (CAT) chooses the next item to locate a person. Inverted CAT runs that logic the other way. It chooses which already-calibrated responses to collect so that the new item is located on the scale, with a standard error, before it is used to measure anything. A seed is a starting guess. The locating step is what turns the guess into a position.

Why this matters for AI safety and AI quality teams

If you sign off on a release, contamination is a defense problem. The question an auditor asks is not only "what did the model score" but "could the model have seen the test." With a fixed benchmark, the honest answer is often "we do not know."

Generating items at measurement time changes that answer. The evaluation set did not exist until the measurement ran. And because each item is located before it counts, this month's score and next month's score are designed to sit on the same scale even though no item repeats.

We think of this as a capability that has to keep moving, not a feature that ships once. It has three parts: notice when a bank or a judge has gone stale, regenerate items along the right pathway, then re-locate and re-check before the new items count. A static benchmark does none of the three.

What exists today, and what does not

This section stays in the post on purpose. Everything above describes a design. This is where it stands.

ComponentStatus (early evidence; not yet certified)
Inverted-CAT locator codeBuilt and tested in early runs. Not yet run on a live item bank with credentialed judges.
Hierarchical pathway (MHC seed)Method designed and certified in design. One methods question about a related use of the seed is open and under review.
Cumulative pathway (S-SBS seed)Code built and unit-tested; design certified. Its within-level validity study is specified and has not yet run.
Credentialed judges to run the pipelineIn progress. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way. None yet meets the full procedure this pipeline requires.
End-to-end run, fresh item to published measureNot yet. Roadmap.

Stated plainly: the architecture and the code are real, and they are further along than a slide. "Contamination-proof, in production" is not true yet, and we will not say it until an end-to-end run with credentialed judges supports it. The next milestones, in order, are the cumulative-pathway validity study, judge credentialing, and a first end-to-end run.

The judge requirement is not incidental. A fresh item located by an uncalibrated judge inherits that judge's lean. That is why the companion post, Evaluating the evaluator, is the other half of this one.

Frequently asked questions

What is benchmark contamination?

Benchmark contamination happens when evaluation items, or close copies of them, appear in a model's training data. The score then reflects how well the model remembers the items rather than how well it can do the task the benchmark was built to measure.

Can an LLM evaluation be made contamination-proof?

Refreshing a benchmark by hand delays contamination but does not end it, and each hand-written set is a new instrument with an unknown calibration. The durable approach is to generate new items at measurement time and place each one on a calibrated difficulty scale, with a standard error, before it counts. XLNC's AIM pipeline is built for this; it is early evidence, not yet certified, and is not running in production today.

What is inverted computerized adaptive testing (inverted CAT)?

Ordinary computerized adaptive testing chooses the next item to locate a person on a scale. Inverted CAT runs the logic the other way: it chooses which already-calibrated responses to collect so that a new item can be located on the scale, with a stated standard error, before that item is used to score anyone.

Is AIM's fresh-item generation available now?

Not as a production service. The locator code is built and tested in early runs, the two seed pathways are designed, and the judges needed to run the pipeline end to end are still being credentialed. XLNC offers early access conversations for teams who want to follow or shape the pilot.

Where this goes next

If your team is deciding releases on a benchmark you suspect has leaked, we would like to hear how you handle it now. Early access is by conversation, not self-serve: join the waitlist and tell us what you measure.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.