The Network That Learned Like a Child
Before a score gates a release or a hire, ask for its papers: the reference it was compared with, the chain back to that reference, and the uncertainty at the end.

In 1993 a cognitive scientist trained simple recurrent networks on a toy grammar. They learned it only when they began with limited memory or simplified input and then matured (Elman, 1993). Start small, then grow: it read like a lesson from child development, written in code.
Six years later, two other researchers ran their own simulations. Starting small, they reported, was not necessary and could even hinder learning (Rohde & Plaut, 1999).
Here is the puzzle. A neural network is among the most re-runnable subjects in science. It does not tire, move away or remember the last session. If careful teams can reach opposite conclusions about one, what chance do findings about people have? And what would it take for a third lab, one that was in neither room, to check either result?
Two fields, one decade of doubt
Psychology found out first. The Open Science Collaboration (2015) repeated 100 published studies; 36% of the replications had statistically significant results. Many Labs 2 tested 28 findings across 125 samples (Klein et al., 2018); 15 reached conventional significance, 14 under the authors' stricter criterion.
Computer science followed. Kapoor and Narayanan (2023) reported data leakage affecting 294 papers across 17 fields; in one civil-war prediction case, the advantage claimed for machine learning disappeared once the leakage was corrected. A review of 445 language-model benchmarks found widespread weaknesses in construct validity (Bean et al., 2025).
The two crises differ in kind. Psychology's headline is whether published effects replicate; machine learning's is partly whether data, code and benchmarks hold up. Yet one shared root, though not the only one, is measurement: constructs undefined, uncertainty unstated, procedures undocumented (Flake & Fried, 2020).
That is the first clue. A finding that cannot say what it was measured against cannot travel to the next lab, and a disagreement between two labs cannot be settled.
A score needs a chain of custody
Evidence counts in court only if you can show who held it and how. In Massachusetts that rule reached measurement directly. A state office certifying breath-test machines withheld records of calibration failures. In April 2023 the Supreme Judicial Court found that about 27,000 defendants' due-process rights had been violated, and breath-test results from June 2011 to April 2019 were excluded; each defendant must seek a new trial (Scott & Andersen, 2023). The ruling turned on the withheld records, not on any single reading.
Metrology gave this idea its international formal shape in 1993. The International Vocabulary of Metrology defined traceability as a property of a measurement result: it can be related to stated references through an unbroken chain of comparisons, each with a stated uncertainty (BIPM et al., 1993a; Belanger et al., 2001). The same year, the Guide to the Expression of Uncertainty in Measurement standardized how to state that uncertainty (BIPM et al., 1993b). The habit is far older; by about 2700 BCE Egypt had made the royal cubit its official standard for central-government building (Hirsch, 2013). In 1999 Mars Climate Orbiter was lost after ground software reported thruster impulse in pound-force seconds where the interface specification required newton-seconds (Mars Climate Orbiter Mishap Investigation Board, 1999).
Now apply the frame to model and human scores. For a model's score, custody means you can say which model answered, which items it saw, what scored it, and with how much uncertainty. For a person's score the questions are the same, with a rater and a form in place of a judge model. Repeatability is not a reference: a judge model that returns the same verdict every time can be wrong every time.
Why the edge of learning needs a ruler
Back to 1993. Elman's networks came from a biological idea. Neural networks began as a model of neurons (McCulloch & Pitts, 1943; Rosenblatt, 1958), so it is no surprise that a lesson from human development was tried on them. The analogy has limits: matching a detailed simulation of one cortical neuron took a deep network of 5 to 8 layers (Beniaguev et al., 2021). The shared lineage explains why the analogy is tempting, not why it holds.
The developmental idea has a name. Vygotsky's zone of proximal development is the gap between what a learner does alone and what the learner does with guidance (Vygotsky, 1978). A meta-analysis of 144 studies of computer-based scaffolding, a practice derived from the zone, found a medium effect (g = 0.46; Belland et al., 2017). Cui and Sachan (2025) went a step further: they used an item response model to locate a language model's zone and to choose when worked examples help.
That step is the second clue. Starting small presumes you know where small is. The zone has to be located before anyone teaches to it, and one measure shows only unaided performance. So "alone" and "with help" must sit on one scale, each with its uncertainty. The realistic target is the next step: coaching for a person, worked examples or fine-tuning for a deployed model. Then measure again, on the same ruler, so the gain is a measured change and not two numbers from two instruments.
What we are building
Our canonical instrument uses Commons' Model of Hierarchical Complexity, which orders completed tasks: an action at each higher order coordinates two or more actions of the order below in a non-arbitrary way. It is designed to become a calibrated Rasch ruler. Once calibration on human data is complete, the instrument will report every measure with its standard error, for leaders and models alike. Until then it is frozen, not ready for customer use; synthetic results show internal consistency only, never superiority.
The same discipline holds for ordered-level constructs such as HEXACO Honesty-Humility and Hallucination, built with Mark Wilson's construct maps inside the measurement hexagon of Luca Mari, Mark Wilson and Andrew Maul (Mari et al., 2023; Wilson, 2005). Our HEXACO work uses published, human-written items. Our Hallucination items and keys are written by language models (GLM, Grok and DeepSeek), checked by an evidence-presence gate and a truth spot-audit. The public reference those values should trace to does not exist yet; a certified referee is pending. That is a stated, open link in the chain. These builds are early evidence, not yet certified.
Lectica's computer-scored instruments are standardized, grounded in Fischer's skill theory, and publish reliability evidence (Lectica, n.d.-a, n.d.-b). Aiden Thornton, who co-authors MASEMS with XLNC's Matt Barney, studied the gap between leaders' reasoning and the complexity their roles demand (Thornton, 2023) and has developed a standardized complexity-leadership instrument. Our path differs by design in theory base (Commons), scoring (a standard error with every measure) and scope (humans and AI on one scale).
NIST AI 800-3 (Keller et al., 2026) names Rasch and item response theory in its case for latent-capability models, whose intervals keep their meaning whichever tasks a benchmark uses. That property is what lets a measure travel down a chain. AIM implements at measurement grade the latent-trait/GLMM approach that NIST AI 800-3 highlights as a promising foundation for AI evaluation statistics.
Ask for the papers
Before a score gates a release or a hire, ask for its papers: the reference it was compared with, the chain back to that reference, and the uncertainty at the end. A score that has them can settle a disagreement between two labs. A score that lacks them gets defended from memory, or rerun at your cost.
We are learning where these chains break across the measurement and evaluation of people and models. Send us a score you rely on and tell us what decision rides on it. Reply on X to @XLNC_CO or use the form at xlnc.co.
Start small, if you like. Just write down where you started, and what you measured it with.
One Ruler. Every Mind.
References (APA 7)
- Bean, A. M., Kearns, R. O., Romanou, A., Hafner, F. S., Mayne, H., Batzner, J., Foroutan, N., Schmitz, C., Korgul, K., Batra, H., Deb, O., Beharry, E., Emde, C., Foster, T., Gausen, A., Grandury, M., Han, S., Hofmann, V., Ibrahim, L., . . . Mahdi, A. (2025). Measuring what matters: Construct validity in large language model benchmarks. arXiv. https://arxiv.org/abs/2511.04703 (NeurIPS 2025 Datasets and Benchmarks Track)
- Belanger, B., Rasberry, S., Garner, E., Brickencamp, C., & Ehrlich, C. (2001). Traceability: An evolving concept. In D. R. Lide (Ed.), A century of excellence in measurements, standards, and technology: A chronicle of selected NBS/NIST publications, 1901-2000 (NIST Special Publication 958, pp. 167-171). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.958
- Belland, B. R., Walker, A. E., Kim, N. J., & Lefler, M. (2017). Synthesizing results from empirical research on computer-based scaffolding in STEM education: A meta-analysis. Review of Educational Research, 87(2), 309-344. https://doi.org/10.3102/0034654316670999
- Beniaguev, D., Segev, I., & London, M. (2021). Single cortical neurons as deep artificial neural networks. Neuron, 109(17), 2727-2739.e3. https://doi.org/10.1016/j.neuron.2021.07.002
- BIPM, IEC, IFCC, ISO, IUPAC, IUPAP, & OIML. (1993a). International vocabulary of basic and general terms in metrology (2nd ed.; ISO Guide 99:1993). International Organization for Standardization. Current edition: JCGM 200:2012, https://www.bipm.org/en/committees/jc/jcgm/publications
- BIPM, IEC, IFCC, ISO, IUPAC, IUPAP, & OIML. (1993b). Guide to the expression of uncertainty in measurement (1st ed.). International Organization for Standardization. Current edition: JCGM 100:2008, https://www.bipm.org/en/committees/jc/jcgm/publications
- Cui, P., & Sachan, M. (2025). Investigating the zone of proximal development of language models for in-context learning. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 6470-6483). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.362
- Elman, J. L. (1993). Learning and development in neural networks: The importance of starting small. Cognition, 48(1), 71-99. https://doi.org/10.1016/0010-0277(93)90058-4
- Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456-465. https://doi.org/10.1177/2515245920952393
- Hirsch, A. P. (2013). Ancient Egyptian cubits: Origin and evolution [Doctoral dissertation, University of Toronto]. TSpace. https://utoronto.scholaris.ca/server/api/core/bitstreams/5d624136-645c-4800-bdae-b19d3117973f/content
- Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804
- Keller, D., Kwegyir-Aggrey, K., Steed, R., Rao, A. K., Sharp, J. L., & Bergman, A. S. (2026). Expanding the AI evaluation toolbox with statistical models (NIST AI 800-3). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-3
- Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Jr., Alper, S., Aveyard, M., Axt, J. R., Babalola, M. T., Bahník, Š., Batra, R., Berkics, M., Bernstein, M. J., Berry, D. R., Bialobrzeska, O., Binan, E. D., Bocian, K., Brandt, M. J., Busching, R., . . . Nosek, B. A. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. https://doi.org/10.1177/2515245918810225
- Lectica. (n.d.-a). About Fischer. Retrieved October 1, 2026, from https://lectica.org/about/fischer
- Lectica. (n.d.-b). CLAS. Retrieved October 1, 2026, from https://lectica.org/about/clas
- Mari, L., Wilson, M., & Maul, A. (2023). Measurement across the sciences: Developing a shared concept system for measurement (2nd ed.). Springer. https://doi.org/10.1007/978-3-031-22448-5
- Mars Climate Orbiter Mishap Investigation Board. (1999, November 10). Phase I report. NASA. https://llis.nasa.gov/llis_lib/pdf/1009464main1_0641-mr.pdf
- McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. Bulletin of Mathematical Biophysics, 5(4), 115-133. https://doi.org/10.1007/BF02478259
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716
- Rohde, D. L. T., & Plaut, D. C. (1999). Language acquisition in the absence of explicit negative evidence: How important is starting small? Cognition, 72(1), 67-109. https://doi.org/10.1016/S0010-0277(99)00031-1
- Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386-408. https://doi.org/10.1037/h0042519
- Scott, I., & Andersen, T. (2023, April 26). Nearly eight years of breath test results cannot be used in drunk-driving prosecutions, SJC rules. The Boston Globe. https://www.bostonglobe.com/2023/04/26/metro/years-breathalyzer-results-cannot-be-used-drunk-driving-prosections/
- Thornton, A. (2023). Facing the complexity gap: Developing leaders' reasoning skills to meet the complex task demands of their roles [Doctoral dissertation, The University of Western Australia]. UWA Research Repository. https://doi.org/10.26182/ky2m-3e90
- Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes (M. Cole, V. John-Steiner, S. Scribner, & E. Souberman, Eds.). Harvard University Press. ISBN 978-0-674-57629-2
- Wilson, M. (2005). Constructing measures: An item response modeling approach. Lawrence Erlbaum Associates. ISBN 978-0-8058-4785-7