For transformation leaders and Chief AI Officers
Most AI projects never pay back.
RAND interviewed 65 experienced data scientists and engineers and reports that, by some estimates, more than 80 percent of AI projects fail, about twice the rate of other IT projects (Ryseff, De Bruhl & Newberry, 2024). Two of the root causes it names are solving the wrong problem and handing AI a task it cannot do.

Human-ARROW targets both causes before money moves.
- It starts at the bottleneck. The constraint in your core process sets the output of the whole process, so that is where we look first.
- It checks the performer before the switch. A person, an AI and a robot are measured on one ruler, against the task's standard, before anyone implements anything.
Ryseff, J., De Bruhl, B. F., & Newberry, S. J. (2024). The root causes of failure for artificial intelligence projects and how they can succeed: Avoiding the anti-patterns of AI (RR-A2680-1). RAND Corporation. https://doi.org/10.7249/RRA2680-1
Human-ARROW: the right performer for each task, decided on one ruler.
The figure above is the whole decision. Find the constraint, measure who or what can do each task there, and pick one of nine pathways. Every choice comes with an error bar, so it can be checked and defended.
The frame around the improvement methods your teams already use.
- Decide who or what does each task. Human-ARROW chooses person, AI or robot, task by task inside the bottleneck, from measurement and evaluation.
- Improve the work you keep. Lean Six Sigma improvement in its Industry 5.0 form: human-centric, resilient and sustainable (Breque et al., 2021).
- Redesign the work you change. Lean Six Sigma redesign, with the same Industry 5.0 aims built in from the start.
Your teams keep the methods they already run. Transformation stops being a bet on a technology and becomes a measured change in who does the work, checked the same way your teams already check every process change.
Built by a Six Sigma practitioner: our founder co-wrote The New Six Sigma (with Tom McCarty), led Motorola's Digital Six Sigma practice, and has applied it as consultant and Champion with health systems and global manufacturers. About the founder
For corporate universities and workforce development programs
Can your learners do it on their own?
That is the question an employer asks the day after the course ends. AI tools now make it harder to answer, because a polished answer no longer shows who did the thinking.
By learning and development we mean all of it: formal training, coaching, performance support tools people use on the job, and the practice built into everyday work. Each one is judged the same way, by whether people can then do the work on their own.
A field experiment with about 1,000 high school math students tested this directly (Bastani et al., 2025). Students practicing with an unrestricted GPT-4 chat tool scored 48 percent higher on practice problems. On the later exam, taken without AI, they scored 17 percent lower than students who never had it. A version of the tool built to give hints, not answers, raised practice scores by 127 percent and largely removed the exam loss.
So AI did not decide the outcome. The design around it did. Herman Aguinis makes the same point about teaching: good teaching in the age of AI needs what it always needed, namely shared ownership of the outcome, honest attention to who is actually served, methods with evidence behind them, and explicit redesign rather than assumption (Aguinis, 2026). Our design for training, coaching and performance support rests on four working parts.
- Learning in the flow of work. People learn on the tasks they already do. Their everyday work is scored as it happens, so nobody stops to sit a course to find out where they stand.
- Each learner in their own next step. People grow fastest on work just above what they can do alone, in their zone of proximal development (Vygotsky, 1978). We place each person on the ruler with an error bar, then set practice one step above. Not so easy it bores them, not so hard it breaks them: the Goldilocks zone, for every learner, not the class average.
- Just in time, not just in case. Before a task, the learner gets a short feedforward note: the one move at their next step that this task calls for. Help arrives when the work arrives.
- Coaching that supports performance and career. Coaches and managers see where a person stands, how steady they are from one occasion to the next, and which roles their measured strengths point toward.



How we tell learning from borrowed answers
We measure the same skill two ways on one ruler: with AI allowed, and on short unaided occasions. The gap between the two is reported as a number with its own error bar. A shrinking gap means the skill is moving into the person. A score that holds steady across many unaided occasions means it is theirs.

This is our design. We have not yet published learning outcomes from it, and the exam figures above come from a school study, not a workplace.
Inside Human-ARROW: the right performer for each task
Five targets for each task
For each task inside the bottleneck we set five targets. The first four, quality, cost, quantity and cycle time, come from Barney (2013); consistency over time is our own addition:
- Quality: how well the work is done.
- Cost: what each unit of work costs.
- Quantity: how far the work can scale.
- Cycle time: how fast the work is done, or its latency.
- Consistency over time: whether the performer does it the same way next month.
How we measure, in plain words
- Active: we ask people and machines short questions with Adaptive Intelligent Measurement. Each question is chosen from the answers so far, out of a bank of automatically generated questions, and the test stops once the score is precise enough. This is computer-adaptive testing (CAT).
- Passive: we score the work people and machines already produce and track its quality over time. We call this Inverted CAT: the everyday work supplies the evidence, so nobody stops to take a test.
Built for Industry 5.0
The European Commission describes Industry 5.0 as human-centric, resilient and sustainable (Breque et al., 2021). Human-ARROW makes the human-centric part checkable. People, AI and robots sit on one scale, so every task assignment is a decision you can audit.
What we measure in people, and why it carries over to AI
Decades of work analysis describe what a worker brings to a job. O*NET, the US occupational information network, organizes worker attributes into cognitive abilities, personality and work styles, interests and values, psychomotor and physical abilities, and skills and knowledge (Peterson et al., 1999, 2001). These are the attributes Human-ARROW compares against a task's demands.
Several of these frameworks now apply to machines as well. HEXACO, a six-factor model of personality (Honesty-Humility, Emotionality, Extraversion, Agreeableness, Conscientiousness, Openness to Experience), is well supported in people (Ashton & Lee, 2007). Computer scientists have since administered personality instruments to large language models and found measurable, shapeable trait profiles (Miotto et al., 2022; Serapio-García et al., 2025). So XLNC measures and evaluates people and AI on one ruler, with the same methods.
References
- Aguinis, H., & O’Boyle, E., Jr. (2014). Star performers in twenty-first century organizations. Personnel Psychology, 67(2), 313–350. https://doi.org/10.1111/peps.12054
- Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150–166. https://doi.org/10.1177/1088868306294907
- Barney, M. (2013). Leading value creation: Organizational science, bioinspiration, and the cue see model. Palgrave Macmillan. https://doi.org/10.1057/9781137361509
- Breque, M., De Nul, L., & Petridis, A. (2021). Industry 5.0: Towards a sustainable, human-centric and resilient European industry. Publications Office of the European Union. https://doi.org/10.2777/308407
- Commons, M. L. (2008). Introduction to the model of hierarchical complexity and its relationship to postformal action. World Futures, 64(5–7), 305–320. https://doi.org/10.1080/02604020802301105
- Commons, M. L., Trudeau, E. J., Stein, S. A., Richards, F. A., & Krause, S. R. (1998). Hierarchical complexity of tasks shows the existence of developmental stages. Developmental Review, 18(3), 237–278. https://doi.org/10.1006/drev.1998.0467
- Miotto, M., Rossberg, N., & Kleinberg, B. (2022). Who is GPT-3? An exploration of personality, values and demographics. In Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (pp. 218–227). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.nlpcss-1.24
- O’Boyle, E., Jr., & Aguinis, H. (2012). The best and the rest: Revisiting the norm of normality of individual performance. Personnel Psychology, 65(1), 79–119. https://doi.org/10.1111/j.1744-6570.2011.01239.x
- Peterson, N. G., Mumford, M. D., Borman, W. C., Jeanneret, P. R., & Fleishman, E. A. (Eds.). (1999). An occupational information system for the 21st century: The development of O*NET. American Psychological Association.
- Peterson, N. G., Mumford, M. D., Borman, W. C., Jeanneret, P. R., Fleishman, E. A., Levin, K. Y., Campion, M. A., Mayfield, M. S., Morgeson, F. P., Pearlman, K., Gowing, M. K., Lancaster, A. R., Silver, M. B., & Dye, D. M. (2001). Understanding work using the Occupational Information Network (O*NET): Implications for practice and research. Personnel Psychology, 54(2), 451–492. https://doi.org/10.1111/j.1744-6570.2001.tb00100.x
- Serapio-García, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., & Matarić, M. (2025). A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence, 7(12), 1954–1968. https://doi.org/10.1038/s42256-025-01115-6
Hiring: job analysis first, then three looks at the same dimensions.
Self-report questionnaires are easy to inflate, and unstructured reference calls predict little. Our hiring design starts with a job analysis: the critical incidents of the role set the dimensions that matter and the standard each must clear.
- Structured reference checks. An AI interviewer asks five or more people who know the candidate's work about specific situations. Each reference's leniency is measured and corrected, as we do for AI judges.
- AI interview. New situations that re-test the same job-related dimensions, rather than starting fresh.
- Human interview. Your interviewer probes the dimensions where the earlier looks disagree. The interviewer is measured as a rater too.
Every stage adds evidence to one picture of the person, so the report is a range, not a single point. The headline figure is the share of work situations in which the candidate clears every standard the job analysis set, reported with its error bar and the number of sources behind it.
This is our design. No hiring results have been published yet.
Coaching: one ruler for coach and client.
A coach needs to know where a client stands, how much the client varies from one situation to the next, and whether the work is helping. Repeated measurement on one ruler, with error bars, gives all three. Coaching is the Human-Grow pathway of Human-ARROW: the move that keeps the work with the person and builds capability.
No coaching results have been published yet.
What is proven today.
- The reasoning ruler passed its gate on 2026-08-22: a standard error at or below 0.07 of a level, at every level from 8 to 12.
- On human personality data, the HEXACO six-factor structure held in a study filed on OSF before the run: all six factors supported.
- In a simulation on HEXACO and IPIP data, the adaptive test reached the precision of the best five-item short form with a median of two to three items.
Customer outcomes from transformation programs are not published yet.
Talk to us about a transformation program About the founder
Aguinis, H. (2026, September 24). Teaching in the age of AI [Article]. LinkedIn. https://www.linkedin.com/pulse/teaching-age-ai-herman-aguinis-kdl4e/
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.