What you pay for
One Coarse Judge
Measure first. Pay when it must hold up.
Start with the cheapest judge that can. Pay for a dated report when a decision has to survive scrutiny.
Plans.
Counts below are measures a month. Each measure is one model on one scale, built from many graded answers, and it ends with its error bar.
Scoped with your team.
Enterprise license
Talk to us.
On your side, scoped with your team
- The calibrated harness runs inside your environment. Your data stays there.
- Your models measured on one ruler, each result with its error bar.
- Dated reports that hold up when a decision is reviewed.
- Support with a set response time.
Proof engagement
Talk to us.
One decision, scoped with your team
- We measure the models behind one decision you face.
- On your hardest case, not a benchmark.
- A dated report with error bars.
- A readout with your team.
Self-serve, by request.
Measurement Program
$1,999 a month
or $19,990 a year. 2,000 measures a month, full judge panel.
- Each measure is a full adaptive run, not one graded answer.
- Starts on the cheapest capable judge. You set the stops.
- Automatic routing to a stronger judge is in build, included when released.
- Top-up packs from $10 when a month runs over.
Start here
Multiple Precise Judges
$499 a month
400 measures a month, full judge panel
- Each measure is a full adaptive run, not one graded answer.
- The full panel gives the tighter error bar.
- Starts on the cheapest capable judge. You set the stops.
- Automatic routing is in build, included when released.
One Coarse Judge
$199 a month
400 measures a month, one judge
- One judge, graded reliably at levels six to ten.
- Above ten, its scores are not counted.
- The lower price buys a wider error bar, not the same answer.
Free
$0
10 measures a month, by request
- Screening precision, so a wider error bar on every measure.
- Each measure is a full adaptive run.
- Not a free run on your own data.
Add-ons: top-up packs from $10. Judge calibration from $2,500 (details). Self-hosted harness, $250 a month, on your side with your own model endpoints.
Hosted plans open by request. There is no checkout yet.
Start cheap, escalate on evidence.
Cheapest judge first. A stronger judge on stand-by, in build. Stops you set: the widest error bar, the most questions, the most spend.
In build.
| Judge | Per 1,336 ratings | L5 | L6 | L7 | L8 | L9 | L10 | L11 | L12 |
|---|---|---|---|---|---|---|---|---|---|
| Jev | US$0.00 | not admitted | admitted, whole level only | admitted, whole level only | admitted | admitted, whole level only | admitted, whole level only | not admitted | not admitted |
| openai-gpt-oss-120b | US$0.50 | not admitted | not admitted | admitted, whole level only | not admitted | admitted, whole level only | not admitted | not admitted | not admitted |
| qwen3-235b-a22b-instruct-2507 | US$1.16 | not admitted | admitted | admitted, whole level only | admitted, whole level only | admitted | admitted, whole level only | not admitted | not admitted |
| deepseek-v4-1-flash | US$4.41 | not admitted | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted | not admitted | not admitted |
| kimi-k3 | US$29.75 | not admitted | not admitted | admitted, whole level only | admitted, whole level only | admitted, whole level only | admitted, whole level only | not admitted | not admitted |
Calibrate your own judge.
One judge, one model, one scale. We measure how harsh or lenient the judge is at each level, correct it onto the ruler, and hand you the dated report.
- Judge calibration, one model, one scale, levels 6 to 10: from $2,500.
- Each added scale for the same model: $1,000.
- Recalibration when your model version changes: $1,000.
- Levels 11 to 13: by request.
You supply the model endpoint and pay for its model calls.
Recertification.
Recertification is part of the cost model. It runs when a trigger fires, and the cost is shared across all tenants. It is not billed as a surprise.
What numbers must you defend?
You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.