XLNCXLNC Watch it measure

Published measures

Reasoning ruler, error bar at each level

LevelStandard errorJudgments
Level 8 (Concrete)± 0.070342
Level 9 (Abstract)± 0.056230
Level 10 (Pre-Formal / Abstract-Formal transition)± 0.056230
Level 11 (Formal)± 0.032232
Level 12 (Systematic)± 0.019225

No total row. Gate passed 2026-08-22. Unit: one level of the ruler.

Trace

Each row resolves to its question IDs, their answer keys and the bank fingerprint, in the dated report of the certifying run.

Every published result, with its date.

The ruler's error bar at each level. Each judge's ceiling, with its pass counts. Never a bare rank.

The precision we publish.

Measured on the reasoning ruler. Standard error 0.07 of a level or better at levels 8 to 12, gate passed 2026-08-22. Newer scales must clear 0.10 at each level, and 0.15 at level 1.

Has the ruler moved?

The question set is frozen, fingerprinted and dated. Every run checks the fingerprint first. Whether the ruler holds steady over months is the next exhibit, and it is not done yet.

Elsewhere, models have moved. Same model names, three months apart, gave different results.

Our re-measurement: not yet shown.

Re-measurement run: not yet published.

Four panels, March versus June 2023, whole percent. Prime versus composite accuracy: GPT-4 84 to 51, GPT-3.5 50 to 76. Counting happy numbers: GPT-4 84 to 35, GPT-3.5 31 to 48. Sensitive questions answered: GPT-4 21 to 5, GPT-3.5 2 to 8. Code directly executable: GPT-4 52 to 10, GPT-3.5 22 to 2.
An outside example. Same model names, three months apart: measured behavior moved. Source: Chen, Zaharia and Zou (2024), Harvard Data Science Review 6(2). Values as printed in the paper, rounded to whole percent.

Is the ruler circular?

Each level is fixed before any answer is graded. Three more guards keep the ruler from grading itself.

Levels 7 to 9: a certified set of 12 judges, from different vendors and countries, reads the same texts. Each judge is checked against the others. The credentialing run is in progress, with no results yet.

Levels 10 to 13: the questions are new. One company writes them. A different company attacks them. Neither judges them. The questions stay sealed and unpublished until the judges read them. The judges for these levels are candidates, still being screened.

So the levels are set first. The questions are unseen. And for the new questions, the writer, the attacker and the judges work for separate companies.

A check against outside human experts: still ahead.

See the roster

Judges, levels 7 to 9

In the certified set. Credentialing run in progress, no results yet.

  • Gemma 4 31B Google
  • Gemma 4 26B Google
  • Gemma 3 27B Google
  • GPT-6 Luna OpenAI
  • GPT-OSS 120B OpenAI
  • Mercury 2.5 Inception
  • Mistral Small 3.2 Mistral
  • Nemotron 3 Nano 30B NVIDIA
  • MiMo V2.5 Xiaomi
  • Qwen 3 8 Flash Alibaba level 8 only
  • Qwen3 235B Alibaba level 8 only
  • Jev TypeSafe

Write and attack the questions

  • Writes and repairs the questions Anthropic
  • Attacks the questions xAI if it passes a reliability check; Anthropic as fallback

Judges, levels 10 to 13

Candidates being screened. Not a final roster.

  • A subset of the level 7 to 9 judges levels 10 to 12
  • DeepSeek V4.1 Flash DeepSeek levels 10 to 13
  • GLM 5.3 Flash Z.ai levels 10 to 12
  • Kimi K3 Moonshot level 12 and up
  • GPT-6 Astra OpenAI level 13, one-level trial first

Roster as of 2026-10-02. No Anthropic model judges at any level. No xAI model judges at levels 10 to 13. Flags show each vendor's home country.

Vendor and model names are used only to identify the models XLNC measures. XLNC is not affiliated with, sponsored by, or endorsed by these companies, and their inclusion does not imply their approval of any result. All names belong to their respective owners.

The board.

Two judge ceilings are published. Both sit at level 11, with a stated condition. Ceiling rule met

Neither of these judges is in the current roster. The record stays here, with its date. The current roster is on the Judges page.

A ceiling is a pass or fail result at each level, so each row shows its pass counts and its date.

The bottom of the range is not mapped yet. The board is regenerated after the next judge runs.

JudgeCeiling levelPassesConditionMeasuredStatus
z-ai-glm-5-3
z.ai
11Break at level 12: 5 of 7.Passes above the break do not count.2026-08-19Not in the current judge roster. Record kept.
deepseek-v4-flash-0731
DeepSeek
11Break at level 12: 4 of 7.Passes above the break do not count.2026-08-19Stale since 2026-09-20: the maker released a new version. Not in the current judge roster. Record kept.

Harsh here, lenient there.

Jev1 sign change beyond the band

+0βˆ’6789101112

openai-gpt-oss-120b1 sign change beyond the band

+0βˆ’6789101112

qwen3-235b-a22b-instruct-2507no sign change beyond the band

+0βˆ’6789101112

deepseek-v4-1-flashno sign change beyond the band

+0βˆ’6789101112

kimi-k3no sign change beyond the band

+0βˆ’6789101112

Above 0 = harsher. Below 0 = more lenient. Pattern only; the exact values stay in the dated report. Shaded band: plus or minus 2 standard errors. Solid dot: admitted at that level; hollow: not admitted. A ring marks a sign change that clears the band on both sides. Level along the bottom. Judge registry of record b5ecea0af663, credential vintages 2026-09-20 and 2026-09-23.

The bend a raw score cannot see.

In a simulation on our own judge telemetry, a run set to a standard error of 0.20 came in at 1.16 with judge harshness left in. With it corrected, 0.21.1

Simulation.

  1. Gate passed 2026-08-22. Criterion met: a standard error at or below 0.07 of a level, with at least 30 judgments, at every level from 8 to 12.
  2. Ceiling rule met. Criterion: a judge's ceiling is the highest level it passes before a break; passes above a break do not count. Certifying runs: the reasoning-ceiling probe runs of 2026-08-19 and 2026-08-22, audited by our quality gate.
  3. The simulation used 15,000 judge decisions, dated 2026-09-09. Method and results are in the whitepaper. Whitepaper