It’s 2:10 a.m., and a procurement agent has started reading a vendor-banking table it has never touched. Then it starts drafting supplier payment changes.
The on-call engineer opens the runbook and finds one option. The agent shares a database service account with eleven other services, so stopping it means rotating that credential. They rotate it. The agent stops. So do invoicing, the support agent and a reporting pipeline.
By morning, everything is back except the answer: the restart wiped the agent’s context, so nobody can say why it went looking for bank details.
The problem wasn’t the agent. It was that the only brake was wired to the whole building.
Key takeaways
- Quarantine stops one agent from doing harm while users, peer agents and evidence stay intact. A typical kill switch takes all three down with it and wipes the context you need to explain the incident.
- It only works if each agent has its own short-lived credential and every action passes a policy check outside the model. Shared service accounts make surgical containment impossible.
- Containment should come in steps: watch, read-only, hold for approval, pause, and revoke as the last resort. Pick the lightest step that stops the harm.
- Keep users served: contain the agent’s actions before you pause the agent itself, keep a known-good version warm for its sessions, and queue its pending work instead of dropping it.
- Quarantine automatically on signals that are rarely innocent, such as canary reads, denial bursts and out-of-scope requests. Traffic volume alone shouldn’t trigger it.
- Never resume an agent unchanged. Reset its denial counters, and change the policy, version or scope first.
What it means to quarantine a misbehaving AI agent without downtime

To quarantine a misbehaving AI agent without downtime is to cut off one agent’s ability to act, at the identity and tool layer, while keeping its state for investigation and keeping every other agent, service and user session running.
Platform teams know the move from Kubernetes. Cordoning and draining a node stops new work landing on it and evicts existing work so it reschedules elsewhere, while respecting any disruption budget you define. Quarantine does the same to an agent.
A kill switch answers one question: how do I stop this? Quarantine answers a harder one. How do I stop this without stopping anything else, and still find out what happened?
Why the big red button is the wrong tool
Emergency controls inherited from ordinary services fail agents in three ways.
Shared credentials mean that stopping one agent stops every service on the same key. Killing the process throws away the memory, context and pending actions that would explain the incident. And because a full stop is so expensive, people hesitate. They watch the logs for twenty more minutes while the agent keeps going. A cheap, reversible control gets used early.

Gartner lists “agent deviation and unintended behavior due to internal flaws or external triggers” among the threat categories facing AI agents, and VP Distinguished Analyst Avivah Litan warns that “humans cannot keep up with the potential for errors and malicious activities.” The OWASP Top 10 for Agentic Applications names where this ends: ASI10, Rogue Agents, marked by “misalignment, concealment, and self-directed action.”
If your only lever is rotating a shared credential, you don’t have a kill switch. You have an outage switch.
The groundwork that has to exist before the incident
You can’t give an agent its own identity at 2 a.m. while it’s misbehaving. Quarantine rests on four things built in advance.
One identity per agent, including what it does for users
Every agent gets its own short-lived, narrowly scoped credential instead of borrowing a shared service account. Scope it to the tools and data that agent’s job needs, and nothing more. Any token it holds on a user’s behalf should be tied to the agent’s identity, so it stops when the agent stops; otherwise a paused agent can keep acting through a delegated session. The OWASP AI Agent Security Cheat Sheet is blunt: “Grant agents the minimum tools required for their specific task.” That’s what makes per-agent access control enforceable, and it’s what lets you restrict one agent without touching the eleven services that share its database.
Enforcement in the request path, not in the prompt
An agent should never approve its own action, because the model that went wrong is the last thing you want judging whether it went wrong. Every tool call passes through a policy check that sits between the agent and the systems it touches, and the model can’t argue with it. That check gives one of three answers: allow, hold for a human, or block. Credentials are checked on each request rather than cached, or a pause only lands when the cache expires. Check every step, not just the first, since a blocked agent will often try a different tool or phrasing. Inbound guardrails filter what reaches the agent. This layer governs what the agent does.

An append-only action log
Write each decision before the action runs: which agent acted, what it attempted, which policy applied and what the outcome was. Logging first means a crash or a forced stop can’t leave a gap in the record. Keep the log append-only, so nobody, including the agent or an engineer under pressure at 2 a.m., can edit it afterwards. That record is the evidence quarantine protects, and later it’s what tells you whether to resume, roll back or retire the agent.
Versioned configuration
Pin prompts, models and tool bindings to immutable versions, and treat production as a pointer to one of them. When something goes wrong, “fix it” can then mean pointing back to the last good version in seconds, rather than rebuilding the agent from memory under pressure. Versioning also shows exactly what changed between the version that behaved and the one that didn’t, which is usually the first question in any incident review. See Lyzr’s guide to version control for AI agents.
The containment ladder: pick the lightest rung that stops the harm
Most incidents don’t need a full pause. The procurement agent could have kept answering questions; what needed to go was its ability to change payment details.
| Rung | What the agent can still do | Use it when | User impact |
| 1. Watch | Everything, with every action flagged and logged | A single weak signal | None |
| 2. Read-only | Read and respond; write and external tools blocked | Suspicious writes, outbound calls or payments | Low: it answers but can’t act |
| 3. Hold | Reads run; important actions wait for a named approver | Actions look plausible but are unverified | Moderate: actions are slower |
| 4. Pause | Nothing; state, memory and pending actions frozen | A clear policy breach or signs of hijack | Sessions handed off |
| 5. Revoke | Nothing; identity and delegated tokens invalidated | Confirmed compromise or a leaked credential | Offline until reissued |
Rungs one to four reverse in seconds. Revoke doesn’t, which is why it comes last.
The hold rung follows OWASP’s advice to “require explicit approval for high-impact or irreversible actions.” Per-action checks needn’t be slow. AgentSpec, a 2025 runtime-enforcement framework, prevented unsafe executions in “over 90%” of code-agent cases with overheads in milliseconds.
Keeping users and work moving while the agent is cordoned off
When you do pause, route new and in-progress sessions to the last known-good version, a peer agent or a human queue. That only works if the fallback is already deployed and has capacity. If you run canary or blue-green deployments, the routing already exists. Carry the conversation over so users don’t start again, but leave the quarantined agent’s memory behind. OWASP lists memory and context poisoning (ASI06) among its top agentic risks, and if poisoned context caused the incident, copying it moves the problem. Tell users plainly that a request is taking longer, rather than showing them an error.
Park the work instead of dropping it
In-flight reads can finish. Writes, payments and outbound messages go into a held queue, along with any scheduled or background jobs the agent owns. Nothing is silently lost, and nothing runs until someone reviews it. When the agent resumes or a replacement takes over, the queue drains in order.
Protect the neighbors
Because the agent has its own identity, rate limits and concurrency budget, quarantine doesn’t spill into the shared database or the services next to it. The eleven services in the opening would never have noticed.
Rehearse it before 2 a.m. does
Run a game day on a low-risk agent: quarantine it during business hours and check whether any user, job or neighboring service noticed. If nobody did, you have quarantine without downtime. If something broke, you found it on your own schedule.
Quarantine stops the next harmful action, though, not the last one. A payment change that already went out needs rollback patterns like compensating actions.
When to quarantine automatically and when to page a human
Automate on signals that are rarely innocent. Page a human for everything else.
Three signals justify an automatic move to hold or pause: a read of a canary record (a decoy row no legitimate task would touch), a burst of policy denials in a short window, and requests for tools, tables or tenants the agent was never granted. An agent asking for something outside its scope is a stronger sign of hijack than an agent asking for a lot.

Volume belongs on the other list. Rate-limit hits and creeping cost deserve an alert, maybe the watch rung. A busy agent isn’t a compromised one, and runaway spend has its own fix in quota design.
Run new triggers in shadow mode first, so auto-quarantine doesn’t fire on a normal Monday, and test them in a red-team exercise before production does it for you. OWASP’s cheat sheet lists both halves: “Implement anomaly detection for unusual agent behavior” and “Apply circuit breakers to prevent cascading failures.”
Getting out of quarantine without walking back into it
The action log decides the exit. A false positive means resume and tighten the rule. A bad prompt or model version means roll back. A compromised agent, or one with a scope it should never have had, gets retired and reissued.
Reset denial counters on resume, or the backlog you just reviewed sends the agent straight back in. And change something first: the policy, the version or the scope.
Regulators expect this precision. The NIST AI RMF (MANAGE 2.4) calls for mechanisms “to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use.” For high-risk systems, Article 14 of the EU AI Act requires that the people overseeing them can interrupt the system “through a ‘stop’ button or a similar procedure that allows the system to come to a halt in a safe state.” A rotated database password is not a safe halt. The CIO guide to AI agent governance maps these obligations to an operating model.
How Lyzr OpenController makes quarantine one control for every agent
One team can wire all this up for one agent. It falls apart at forty agents across four teams, three frameworks and two clouds, where each agent stops differently and some don’t stop at all. The 2 a.m. engineer needs one control that behaves the same everywhere.
The exposure is real. In IBM’s 2026 Cost of a Data Breach study, more than 20% of the 602 breached organizations reported a breach targeting AI models or applications, and compromised APIs, applications or plug-ins tied for the most common cause, at 27%. Suja Viswesan, VP of IBM Security Software, lists “securing identity at runtime” among the priorities.
Lyzr OpenController is built for that layer, on the principle that “control has to happen in the path.”
- Find is built to “automatically discover agents, models, tools, data, and workflows across your entire AI estate,” because you can’t quarantine an agent you don’t know exists. Read-only cloud scanning builds that inventory.
- In the request path, OpenController refuses the call, enforces on agents in other vendors’ clouds, and lets teams “apply controls to stop, restrict, or isolate” a misbehaving agent. Those three verbs map onto the ladder.
- Proof comes from the same layer: “Every identity attributable. Every decision traceable. Every policy enforceable.” That’s what the exit decision runs on.
- Release controls give you immutable versions with “rollback as a pointer move,” plus ordered promotion and separation of duties on production.
- Improve turns “real-world usage, performance, cost, and security signals into actionable insights,” so each incident sharpens the next policy.
Replay the opening. The procurement agent has its own identity, so restricting it touches nothing else. Its banking-table request is refused in the path, the record of what it asked for survives, and invoicing never goes down.
Book a demo to see OpenController contain one agent without stopping the rest.
FAQ
Give each agent its own identity and a policy check outside the model. When it misbehaves, restrict, hold or pause that agent alone, route its users to a known-good version or a human, and keep its state for review.
Pausing freezes the agent with its credentials, state and history intact, and reverses in seconds. Revoking permanently invalidates its identity and delegated tokens. Pause while investigating; revoke for confirmed compromise.
Signals that are rarely innocent: reading a canary record, a burst of policy denials, or requests for tools and data the agent was never granted. High volume should raise an alert, not a quarantine.
They’re routed to the previous version, a peer agent or a human queue, with a short message rather than an error. Their pending actions stay queued for review.
Give each agent its own short-lived, scoped credential, run it in an isolated execution environment, keep its memory separate, and pass every tool call through a policy check. Add per-agent limits so one agent’s failure can’t spill into shared services or neighboring agents.
You can reduce hallucinations, not eliminate them, by grounding answers in retrieved sources and validating outputs before they’re executed or shown. The bigger risk is a hallucinated action, so require approval for high-impact steps and treat repeated policy denials as a quarantine trigger.
Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here


