XLNCXLNC Watch it measure

Next

Next: robots on the same ruler.

Nothing is measured yet. No robot has a published measure from us. To measure robots we first need to capture video and sensor streams and monitor what they do. Today we support text only. This page explains how the ruler extends, and we will not show a number until one exists.

A person and an AI each shown as a solid distribution on one shared capability ruler; a robot arm and its distribution are drawn dashed and marked Next.
Illustrative schematic. No robot has been measured yet.

Why the ruler does not care what the performer is.

A calibrated ruler places questions and performers on one scale. The performer can be a person writing an answer, a language model producing text, or a robot completing a task. What changes is the evidence: words, speech, pictures, video, or a record of actions in the world. The scale, the error bars and the judge corrections stay the same.

Text, audio, image, video, sensor streams and robot telemetry flow into a calibrated judge panel, which produces one measure with an error bar. Only text is marked Today; the rest are dashed and marked Next.
Illustrative schematic. Only text evidence is graded today.
Five example robot tasks placed as rising steps on the Primary, Concrete, Abstract, Formal and Systematic stages of hierarchical complexity.
Illustrative placements. No robot has been measured on any stage.

What has to be built first.

  1. Video and sensor capture. Prerequisite for everything else: we must record video and sensor streams and monitor robots as they work. Today we support text only. Planned.
  2. Routing by evidence type. The engine already sends non-text material to models that can read it. That routing is built and tested.
  3. Choosing the judges that can see. Picking and qualifying models that can grade images, audio and video, each from a different company than the one that produced the work. Planned.
  4. How much evidence is enough. For each kind of evidence, the smallest amount that reaches the precision we promise. The text version can run on data we already hold; audio, image and video wait on the judges above.
  5. Tracing where the evidence came from. Extending our measurement trail so that every score from a sensor or a recording records its source, the same way a text score records its question and its judge.
A dashed band shows one robot's expected capability across six situations, such as dim light and a slippery surface, above a dashed minimum requirement line.
Illustrative schematic. Shapes show the idea, not data.

Planned work. Nothing on this page is a measured result.