XLNCXLNC Watch it measure

Blog

Calibrated judge scores in your OpenTelemetry traces: what we are building

AIM is building a way to read your OpenTelemetry traces and write calibrated judge scores, with standard errors, back onto your own spans. Trace ingest is in development and offered as early access. It complements your observability stack; it does not replace it.

If your team runs an LLM application in production, you probably already have its behavior on record. Every request, every retrieval, every tool call, every model version, captured as OpenTelemetry spans and sitting in the observability tool you chose.

That record is the most honest evidence of what your system does. What it usually lacks is a trustworthy judgment of whether the output was any good.

Short answer. XLNC is building a way for AIM to read your exported OpenTelemetry traces, score the outputs with calibrated judges, and write each score, with its standard error, back onto your own spans. Trace ingest is in development and offered as early access; it is not a released feature. Score write-back comes after ingest. AIM is designed to complement your observability stack, not replace it.

Status first

Because this page describes something still being built, the status goes at the top.

ComponentStatus
Trace ingest specificationIn development. Not yet certified.
File-based ingest of exported OTel tracesIn development. Early access only, handled directly with each team. No live endpoint.
Calibrated scoring of ingested outputsDepends on judge credentialing. Five judges hold credentials at levels 6 to 10 under our earlier procedure; re-credentialing under our stricter new procedure is under way.
Write-back of scores onto your spansRoadmap. Planned after ingest.
Integrations with specific platformsNone shipped. Nothing on this page claims a live integration.

Stated plainly: if you send us traces today, you are joining an early-access conversation, not turning on a feature.

Complementary to your observability stack

This part is settled, so we will say it once and clearly. AIM is designed to sit alongside the tools you already run, such as Arize Phoenix, LangSmith, Braintrust, Datadog, Honeycomb, or Grafana. It does not replace any of them.

Those tools are good at what they do: collecting traces, showing latency and cost, letting you search and debug. Many of them also run evaluations. AIM's job is narrower. It measures, on a calibrated scale, and it hands the measurement back in the format your tools already read. You keep your stack. Your traces stay yours.

Why the trace is the right place for a score

A score that lives in a separate dashboard gets checked when someone remembers to check it. A score attached to the span it describes shows up next to the latency, the model version, and the prompt that produced it. When a release goes wrong, the evidence and the judgment are in one place.

Claim. OpenTelemetry is the common wire format for LLM traces across observability platforms, and its GenAI conventions are still being settled.

Evidence. Arize Phoenix, Galileo, LangSmith, and Braintrust document ingesting OpenTelemetry traces, and Datadog and Grafana consume OpenTelemetry GenAI spans. The OpenTelemetry GenAI semantic conventions (opentelemetry.io/docs/specs/semconv/gen-ai) remain in Development status as of 2026: none of their spans, events, or attributes is marked Stable, and in June 2026 they moved to a dedicated repository so they can change faster than the core stability bar allows.

That second fact shapes how we build. An emerging, not-yet-stable convention means field names can change. We plan to pin a specific convention version, state it, and accept common aliases, so a change upstream does not silently break your scores.

What the score is designed to carry

The attribute is the easy part. Anyone with OpenTelemetry plumbing can write a number onto a span. What matters is the number.

An AIM score is designed to come from a judge that was measured before its scores counted, on a scale where the work and the judge sit together. That is the subject of Evaluating the evaluator. Every score carries its standard error, so a reader of the trace can tell a score that can bear a release decision from one that cannot.

Our draft specification outline proposes attributes in an aim.* namespace on your own span ids, along these lines. These names are proposals and may change before the specification is certified.

Proposed attributeMeaning
aim.scoreThe calibrated score for the output this span produced
aim.score.seThe standard error of that score
aim.judge.variant_idWhich judge produced the score
aim.verifier_versionA content hash of the rules and prompt that produced the grade

Where the pinned GenAI convention supports a standard evaluation form, the plan is to write that too, so scores appear in tools that read it. Every written-back score will be labeled early evidence until the calibration behind it is certified.

Claim. Writing evaluation scores back onto OpenTelemetry spans is uncommon in public tooling today.

Evidence. In a review of public documentation from major evaluation and observability vendors (Arize Phoenix, Galileo, LangSmith, Braintrust, Datadog, Grafana), we did not find one documenting write-back of a judge score onto a span as an attribute in the GenAI wire format. This is absence of evidence from a fast-moving area, not proof.

A receipt for every grade

A score you cannot audit is a score you have to take on faith. The plan is for each AIM grade to carry a receipt: which trace and item it belongs to, which version of the rules produced it, what the judge actually answered, which rule failed if one did, how long it took, what it cost, and whether a fallback judge answered instead of the primary one. The reasoning behind this, and the near-miss that prompted it, is in The grader that never asked how. This is roadmap, built alongside ingest.

Your data, handled as a measurement record

Traces can contain prompts, outputs, and user details. Our draft specification outline sets rules before any of that is scored: content is opt-in, identifiers are hashed, and files are scanned for personal data at ingest. A file that fails the scan is quarantined and receives zero judge calls. Retention and deletion rules are being set with counsel. These are design commitments in a specification outline still under review, not a certified capability.

Frequently asked questions

Can AIM ingest OpenTelemetry traces today?

Not as a released feature. Trace ingest is in development and offered as early access. The first version is planned as file-based ingest of exported OpenTelemetry traces, handled with each early-access team directly, not a live endpoint.

Do I need to replace my observability tool to use AIM scores?

No. AIM is designed to complement the observability stack you already run. Your traces stay in your tools. The plan is for AIM to write calibrated scores back onto your own spans so they appear where you already look.

Are the OpenTelemetry GenAI semantic conventions stable?

Not yet. The OpenTelemetry GenAI semantic conventions are still in Development status, and none of their spans, events, or attributes is marked Stable. Some observability platforms already support parts of them. AIM plans to pin a specific convention version and state it.

What makes an AIM score different from other evaluation scores on a trace?

The attribute is ordinary; the number behind it is not. An AIM score is designed to come from a calibrated judge on a measured scale and to carry its own standard error, so a reader of the trace can see how much weight the score can bear. Scores will be labeled early evidence until the calibration behind them is certified.

Where this connects

The score on the span is only as good as the judge behind it: Evaluating the evaluator. The items that judge scores against should be ones no model has seen: Fresh instruments on demand.

If you already export OpenTelemetry traces and want a calibrated read on what is in them, join the waitlist and tell us which tools you run. Early access is by conversation.

Tell us which number you do not trust.

You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

We store what you send with the few providers our Privacy Policy names, and use it to reply to you. We do not sell it. Email matt@xlnc.co and we delete it. Read our Privacy Policy.