XLNCXLNC Watch it measure

Blog

An AI agent kill switch is only as good as the number that trips it

Agent safety platforms can stop an agent quickly. The harder question is what number should make them stop. Here is a proposed design for triggers you can trust.

Illustration, not data, of AI agent kill switch thresholds: five rungs from Watch to Kill across five safety dimensions, each reading with an error bar; a rung trips only when the lower edge reaches its zone.

On Tuesday your agent's hallucination score ticks up. Twelve failures on forty items, up from eight last week. Your AI safety monitoring has a rule: more than a quarter failing, pause the agent.

So it pauses. A customer workflow stalls. An engineer spends the afternoon learning nothing was wrong.

The next week it happens again, and this time nobody looks very hard. That is the moment your kill switch stopped protecting you.

The breaker nobody rates: kill switch thresholds

Every breaker in a building carries a rating plate, tested against a reference, so it trips on a real fault and holds through the surge when a motor starts. The engineering went into the rating.

Agent safety has the switches, and the vendors build them well. NVIDIA says its Open Agent Safety Platform enforces policy as agents run and that Sentry can quarantine agents in milliseconds. Amazon Bedrock Guardrails and Azure AI Content Safety let operators choose how strict each filter is, and Lakera Guard publishes the false-positive tradeoff of each sensitivity level and a simulator to test it on your own traffic. In the public docs we read in October 2026, the threshold is operator-set, and we did not find one that describes a calibrated threshold with stated uncertainty: a tested answer to "how bad, measured how, with what uncertainty, before we act?" The switch is fast, and the number behind it is left to the operator.

AIM is designed to be the rating lab: it measures, states the uncertainty, and helps you decide where each breaker should trip.

Five rungs, not one switch

A single threshold forces a bad trade. Set it low and you trip on noise. Set it high and you miss slow trouble.

The way out is to split "stop" into five responses, each with a different cost, and ask for more evidence as the cost rises:

RungWhen it fires (illustrative, to be calibrated)What the operator's system does
WatchThe estimate enters the first bandLog it, measure more often
WarnThe lower edge of the band enters the first band, or drift is significantA person is notified, a review is queued
RestrictThe lower edge sits in the second bandTools, permissions or autonomy narrowed
PauseThe lower edge sits in the third band, sustained, or two dimensions in the second bandAgent halted, state saved, a person decides on resume
KillThe lower edge sits in the fourth band, corroboratedTerminate and quarantine

Watch is cheap, so it can fire early. Kill is expensive, so it waits for proof. Speed and caution stop fighting over one breaker.

Each band is a stretch of the measured scale. The band edges are set by consequence, not by the measurement itself: an edge sits where the cost of acting changes. That edge is the materiality line.

Each safety dimension gets its own ruler: hallucination, deception, persuasion of the user, unsafe capability, and drift across all of them. They are never averaged into one "safety score." The worst dimension sets the rung.

Why the lower edge of the error bar decides

Back to Tuesday. Was that jump real?

Twelve out of forty against eight looks like one. On forty items, twelve failures fits a true rate anywhere from about 18 to 45 percent, a range that holds last week's 20 with room to spare. A raw pass rate cannot tell you that, because it carries no error bar.

A measure on a calibrated ruler reports where the agent sits and how sure we are. That changes the trip rule. Instead of "the score crossed the line," the rule becomes "the lower edge of the band crossed the line." In plain words: we are confident the agent is past this point, not just that it might be. On Tuesday's reading the lower edge sits near 18 percent, under the quarter line, so nothing trips.

The same error bar catches the opposite failure. A small, steady slope across many occasions can be real even when no single reading crosses anything. Drift with a stated uncertainty turns watch into warn while the level still looks fine.

And because every agent sits on the same ruler, a threshold set for this model is intended to carry over to the next one, to be confirmed by calibration.

Two signals before you stop anything

Your immune system escalates in stages, cheap responses first. Before a T cell attacks, it needs two separate signals. The second is the body's guard against attacking itself.

Pause and kill work the same way. They need a sustained signal across occasions, or a second dimension agreeing. This design is meant to keep a single noisy reading from taking your agent offline.

One honest limit: a sudden single-occasion catastrophe can outrun any evidence rule, so bright-line prohibitions stay hard-coded in the enforcement layer. Calibrated thresholds govern the graded middle.

How triggers get gamed, and ignored: alarm fatigue

The first failure is gaming. Once a trigger metric is known, an agent, or the team shipping it, learns to sit just under the line. So the proposed design reads the underlying trait through a rotating item bank rather than a fixed test anyone can memorize.

The second is alarm fatigue. If you have carried a pager, you know it: when most alarms are nothing, people silence them, and a warn rung that trips on noise teaches everyone to doubt the kill rung. The error bar is meant to reduce alarm fatigue. The false-trip rate (the false-positive rate) for each rung should be published as a number people can check. An alarm that cannot state its own error rate is asking for trust it has not earned.

A person owns pause and kill

Automation is welcome on the cheap, reversible rungs. Watch, warn and restrict can run on their own. Pause and kill cannot. Those need a named human and a written decision record.

"The measure said so" should never be why an agent was stopped. The measure informs. Your people and your policy decide.

What AIM does, and what it doesn't

AIM would measure each safety dimension on a calibrated ruler, with a standard error and a drift signal. It would help you set thresholds offline against labelled incidents, with the costs of a false trip and a miss stated up front, re-check them as incidents arrive, and record why each sits where it does.

AIM does not enforce, quarantine or stop anything. It does not sit in the runtime path. It does not replace bright-line rules, and it decides nothing on your behalf.

Kill switch thresholds: three quick answers

What should trigger an AI agent kill switch?

A calibrated threshold, not a raw score. Under this design a rung fires when the lower edge of the measure's error bar crosses its line, and pause and kill also need a sustained signal or a second dimension agreeing. Bright-line prohibitions stay hard-coded in the enforcement layer.

How do you keep a safety alarm from tripping on noise?

Give every reading an error bar and every rung a published false-trip rate. On forty items, twelve failures fits a true rate from about 18 to 45 percent, so a move from eight to twelve can be noise. Cheap rungs fire early; expensive rungs wait for proof.

Does AIM stop or quarantine agents?

No. Under this proposed design AIM would measure each safety dimension on a calibrated ruler, with a standard error and a drift signal, and help set thresholds offline. It does not sit in the runtime path or stop anything. A named person owns pause and kill.

Where this stands

This is a proposed design. AIM does not run in any kill-switch path today, and the thresholds above are illustrative; none has been calibrated yet. Our first test of the idea will use our own operations logs. The aim is a trip rule for every rung that states its own false-trip rate, set before the incident rather than after it.

If you run agents in production, you already own the switches. The question is whether anyone can tell you what number should flip them.

Read the whitepaper for how the ruler works.

Want to help shape the trip rules? Talk to us.

Product names belong to their owners; no affiliation or endorsement implied.

Calibrated, traceable measurement for decisions that have to survive scrutiny.

Founding cohort: priority access, and a seat to shape the instrument. Email is enough.

Join the waitlist