XLNCXLNC Watch it measure

WHITE PAPERS

Measurement and evaluation, written down

The method behind the measures, in full: what we did, every parameter, and every limit, for the people who have to defend a decision about an AI system or a person.

Free. Fill in one short form once; every paper here then downloads in one click on this browser.

All white papers

  • METHOD

    Where AI Transformation Fails

    A task-level method for choosing who or what does the work

    Know which tasks to give to people, which to AI, and which to both, before you spend.

    For: Leaders running AI transformation

    14 pages · PDF · Oct 2026

  • SIMULATION STUDY

    The Error Bar You Buy Is the Error Bar You Get

    Check that the error bar a vendor sells is the error bar you actually get.

    For: AI evaluation and risk leaders

    21 pages · PDF · Oct 2026

    What is inside
    • The problem. Judge severity varies across capability levels, and one average correction cannot represent a bias that changes sign. The parameterization is wrong, so no amount of data fixes it.
    • This study. 15,000 simulated certification decisions over the measured severity curve from a judge probe we ran, with the method appendix, every parameter published, and an honest claim tier: model output at proposed prices, never measured field fact.
    • The three-way contrast. No correction delivers a standard error of 1.16 against the 0.20 sold. The omnibus average leaves 1.34 points of leniency on weak low-band work and invents 0.73 to 1.44 points of harshness on strong work, the sign-flip an average cannot see. Capability-level correction delivers 0.21 against the 0.20 spec at zero marginal cost, with zero mean residual in every capability cell.
    • The asymmetry a buyer should care about. A false pass admits a biased judge whose every later measure inherits the bias, a wave-scale contamination event. A false fail costs only the foregone judge-line saving. The risk sits orders of magnitude on the false-pass side, exactly where the average correction leaves 1.34 points of leniency.
    • Traceability as the mechanism. How per-capability-level correction works, and why it makes every certification decision auditable after the fact, line by line.
  • METHOD

    Job Graph Analysis

    Find the constraint before you redesign the job

    Start redesign at the bottleneck, then measure who or what should do each task there.

    For: Leaders redesigning work around AI and job analysts

    16 pages · PDF · Oct 2026

  • ARCHITECTURE

    Calibrate the gate before the gate decides

    How AIM puts every test item on one ruler before it judges a person, a model, or both

    See which parts of AIM you can rely on today and which are still being built.

    For: Evaluation teams and measurement scientists

    10 pages · PDF · Oct 2026

  • METHOD

    Selecting Talent for the AI Era

    One estimate from many angles: how to combine what hiring science already knows with what AI can now measure

    Combine proven hiring evidence with new AI measures into one estimate, and know how far to trust it.

    For: Talent and HR leaders

    16 pages · PDF · Oct 2026

Get the white papers

Fill in this short form once. The PDF opens as soon as you send, and every other paper here then downloads in one click on this browser.

Which paper?

Please don't include personal or sensitive information.

The PDF opens as soon as you send. One download, no drip sequence.

We remember this browser for 12 months so the other papers download without this form. Use a different email, or clear your cookies, to stop.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.