Skip to content

Guide · agent safety

What is an AI agent kill switch?

An AI agent kill switch is a single action — by a person or an automatic trigger — that stops one agent everywhere it can act, immediately, without stopping the others. Done properly it is more than turning a status to “off”: the agent is refused at the point where its requests arrive, its credentials stop working, anything it queued is cancelled, the stop is recorded as evidence, and only an accountable owner can let it back.

Why status alone is not a kill switch

If stopping an agent only changes a flag, every credential it holds keeps working until someone remembers to revoke it, and an action it queued for approval can still be approved afterwards. A kill switch has to do all of it in one act, because under pressure nobody runs a checklist.

What happens when the switch is thrown

In Agent Trust Cloud, containing an agent does these in order, fastest first:

  • Blocks the agent at the gateway, before anything else is written.
  • Revokes its credentials, cancels its pending approvals and closes its open sessions, in one database transaction.
  • Writes a quarantine record signed over what was done.
  • Opens an incident whose dossier starts before the trigger, and notifies the organization’s owners.

Automatic triggers

The switch can also be thrown by the behaviour monitor. Each agent has a baseline of the tools, destinations, data and spend it normally uses; a score above the quarantine threshold the owner set contains the agent automatically, with the reasons written in plain English. Lower rungs raise an alert or hold every action for a person instead.

How fast, measured

The time from trigger to the first refused agent request is measured by the product’s own end-to-end test on every release, and the page on AI agent monitoring publishes the median and 95th percentile from that run. It is a software control over HTTP, not a hardware one, and is reported in milliseconds only because it was measured in milliseconds.

Letting an agent back

Only an organization owner can release a quarantined agent, with a written reason. It comes back restricted, not in production, and with no credentials; a new one is issued deliberately. NIST’s AI RMF asks for exactly this kind of mechanism — to supersede, disengage or deactivate an AI system that behaves outside its intended use (MANAGE 2.4).

Sources