Skip to content

Behaviour baselines · anomaly response · kill switch

AI agent monitoring that can stop an agent

Each agent gets a baseline of what it normally does — its actions, destinations, data, volume, hours and spend — learned from its own decision record. New activity is scored against it, every point comes with a sentence explaining it, and the score drives a response you choose: an alert, a hold on every action until a person approves, or quarantine through the kill switch.

The response ladder

Score reachesWhat happens
30 (default)An anomaly event with the explanation is written to the evidence ledger.
60 (default)An agent-scoped require-approval rule is added: every action waits for a person until someone removes the rule.
85 (default)The kill switch is thrown automatically, with the explanation as the recorded reason.

Thresholds are set per agent, any rung can be switched off, and nothing acts automatically until an owner configures a ladder for that agent. An agent with fewer than 20 recorded decisions is reported as still learning rather than scored — no history is not the same as normal history.

A worked example

An invoice agent’s baseline: 40 decisions over 2 actions, one destination, working 10:00 and 14:00 UTC. Then, one night, three actions:

Score 100/100 → quarantine

  • +15: Used an action it has never used before: export_vendor_list.
  • +25: Sent to a destination never seen before: files.unknown-host.example.
  • +30: Touched data it has never handled: personal data.
  • +10: Active at 02:00 UTC, hours it has never worked before.
  • +35: Spent $19,000.00 in a day, more than twice its highest day of $420.00.

The kill switch, measured

Throwing the switch blocks the agent at the gateway first, then revokes its credentials, cancels its pending approvals and closes its sessions in one transaction, writes a signed quarantine record, opens an incident and notifies the owners. Only an organization owner can release it, and it comes back restricted with no credentials.

Trigger to first refused agent request: 17 ms median, 81 ms at the 95th percentile over 15 runs.

Inside the server, the gateway block is in place in under 0.1 ms (median) after the kill-switch handler starts, before any database work; the rest is the round trip and the agent’s next call.

Measured on 2026-09-28: local test server, Node v24.19.0, PostgreSQL 16 on the same win32 machine, loopback HTTP. Each run registers an agent, issues it a credential and has it call the gateway continuously, then throws the kill switch through the API. Time runs from sending the contain request to the first refused agent request sent after it.

This is a software control: it needs no special hardware and applies wherever your agents call through the gateway. See what an AI agent kill switch should do.