Next
Next: robots on the same ruler.
Nothing is measured yet. No robot has a published measure from us. To measure robots we first need to capture video and sensor streams and monitor what they do. Today we support text only. This page explains how the ruler extends, and we will not show a number until one exists.

Why the ruler does not care what the performer is.
A calibrated ruler places questions and performers on one scale. The performer can be a person writing an answer, a language model producing text, or a robot completing a task. What changes is the evidence: words, speech, pictures, video, or a record of actions in the world. The scale, the error bars and the judge corrections stay the same.


What has to be built first.
- Video and sensor capture. Prerequisite for everything else: we must record video and sensor streams and monitor robots as they work. Today we support text only. Planned.
- Routing by evidence type. The engine already sends non-text material to models that can read it. That routing is built and tested.
- Choosing the judges that can see. Picking and qualifying models that can grade images, audio and video, each from a different company than the one that produced the work. Planned.
- How much evidence is enough. For each kind of evidence, the smallest amount that reaches the precision we promise. The text version can run on data we already hold; audio, image and video wait on the judges above.
- Tracing where the evidence came from. Extending our measurement trail so that every score from a sensor or a recording records its source, the same way a text score records its question and its judge.

Planned work. Nothing on this page is a measured result.
Hear when the first measure lands.
Join the list How Human-ARROW decides between people, AI and robots