XLNCXLNC Watch it measure

Blog

The practice score went up. The exam score went down.

In a field experiment, students who practiced with an unrestricted AI tool scored 17 percent lower on the later exam. Here is why that happens, and how to measure learning instead of borrowed answers.

2026-10-04 · measurement practice

The finding

Practice grades came in 48 percent higher, and exam scores 17 percent lower, than students without AI. Same students, same semester.

The practice was done with an AI tool that answered any math question. The exam was taken without it.

Bastani and colleagues ran this field experiment with about 1,000 high school math students (Bastani et al., 2025). One group practiced with an unrestricted GPT-4 chat tool. Both figures compare them with students who never had the tool.

A second group used a version built like a tutor. It gave hints, not finished answers. Practice grades rose 127 percent, and the exam loss largely went away.

Same model, same school, opposite outcomes. The tool did not decide what was learned. The design around it did.

The pattern has shown up again. In a preprint, a second team studied college math students on the ALEKS platform (Rismanchian et al., 2026). After ChatGPT arrived, on problems AI can solve compared with problems it cannot, the odds of a correct answer fell 25 percent when students were proctored and rose 85 percent when they were not.

Two panels, each on its own scale. Panel A, Bastani et al. (2025): students who practiced with an unrestricted GPT-4 tool scored 48 percent higher on practice and 17 percent lower on the unaided exam than students without AI. Panel B, Rismanchian et al. (2026), a preprint using ALEKS college math placement data: after ChatGPT, on problems AI can solve compared with problems it cannot, the odds of a correct answer fell 25 percent under proctoring and rose 85 percent without it.

Why practice scores mislead

A practice score answers one question: was the work right? It does not answer a second one: who did the thinking? When an AI tool can do the work, those two questions come apart, and the practice score keeps answering only the first.

Most learning and development programs, whether training, coaching or the performance support tools people use on the job, only collect the first kind of score. Course completions, quiz results and polished assignments all look better when AI helps. None of them shows whether the person can now do the task alone, which is what an employer is paying for.

What the fix looks like

Herman Aguinis argues that teaching in the age of AI needs what good teaching always needed: shared ownership of the outcome, honest attention to who is actually served, methods with evidence behind them, and explicit redesign rather than assumption (Aguinis, 2026). Measurement turns the last two into something a program can check. Five working parts follow.

1. Measure in the flow of work. Score the tasks people already do, as they do them. No separate course is needed to find out where someone stands.

2. Place each person on one ruler, with an error bar. Then set practice one step above where they can work alone. This is their zone of proximal development (Vygotsky, 1978): hard enough to stretch them, not so hard they hand it to the AI.

3. Send help just in time. Before a task, a short feedforward note names the one move at the learner's next step that this task calls for.

4. Measure the gap. Score the same skill with AI allowed and on short unaided occasions, on the same ruler. Report the gap as a number with its own error bar. A shrinking gap means the skill is moving into the person.

5. Watch for steadiness. One good result can be luck or help. A score that holds across many unaided occasions is the person's own.

Who this is for

Corporate universities and university workforce development programs both make a promise to someone else: the learner will be able to do the work. Both now face learners with AI tools that can make every assignment look finished. The promise is kept or broken in the unaided score.

The limits

The Bastani study was in a high school math classroom, not a workplace, and we have not yet published learning outcomes from our own design. What we offer today is the measurement and evaluation: one ruler, error bars, and the gap between assisted and unaided work.

Pick one skill your program teaches. Score it once with AI allowed and once unaided. If you want help reading the gap, see https://xlnc.co/transformation.

References

Aguinis, H. (2026, September 24). Teaching in the age of AI [Article]. LinkedIn. https://www.linkedin.com/pulse/teaching-age-ai-herman-aguinis-kdl4e/

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122

Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster completion, less learning: Generative AI reduced study time on math problems and the knowledge they build (arXiv:2605.21629v3) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21629

Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.

Calibrated, traceable measurement for decisions that have to survive scrutiny.

Founding cohort: priority access, and a seat to shape the instrument. Email is enough.

Join the waitlist