Reversible by Design: The Agentic AI Feature to Demand
Reversible AI actions let agentic AI act on infrastructure safely. How to classify, bound, and enforce reversibility, and prove the host recovered.
By SAUTERA
Why reversible AI actions are the enterprise safety control that separates AI agents you can deploy from AI agents you have to babysit.
Why reversibility beats permission
Agentic AI changed the risk calculus. A chatbot suggests; an agent acts. When software can send emails, move money, revoke access, or patch production, the question is no longer "is the answer good?" but "can we undo it if it isn't?"
Most enterprise AI safety programs answer the wrong question. They gate AI agents with permissions — who is allowed to do what — and stop there. That is necessary but not sufficient. Permission tells you an action was authorized. It says nothing about whether the action can be walked back when the model is confidently wrong.
Reversible AI actions flip the default. Instead of trusting the AI agent to be right, you design the system so that being wrong is cheap. Identity tells you who is acting; as we put it, zero trust tells you who, not whether. Reversibility is part of answering whether.
The stakes are concrete. Gartner projects that by 2028, 15% of day-to-day work decisions will be made autonomously by agentic AI, up from essentially zero in 2024. At that volume, an agent that is 99% reliable still produces thousands of wrong actions per week. Reversibility is what keeps those errors from compounding into incidents.
- Permission asks: is this AI agent allowed to act?
- Reversibility asks: if it acts wrongly, what is the blast radius and recovery time?
- Enforcement asks: can we guarantee the answer to both before the action runs?
What reversible by design actually means
Reversibility is not "we keep backups." A nightly snapshot is a hope, not a control. Reversible by design means every AI agent action carries three properties before it executes.
It is classified by consequence. The system knows, at design time, whether an action is trivially reversible (draft a document), reversible with cost (send an internal message), or effectively irreversible (wire funds, delete a customer record, publish to the public). This classification is the foundation. Everything else keys off it.
It has a defined undo path. For every reversible action, there is a specific, tested mechanism to reverse it — not a vague promise. If the undo path does not exist, the action is not reversible, and the system should treat it as high-consequence regardless of how routine it looks.
It is bounded before it runs. Irreversible actions do not execute on the AI agent's word alone. They pause for a human, a second control, or an explicit policy that says "this class of action is permitted here." This is enforcement, not advice. As we argue in continuous, observed, enforced, a control that is not enforced at runtime is decoration.
The payoff is speed. Teams assume safety controls slow AI agents down. The opposite holds when reversibility is designed in. When the blast radius of most actions is near zero, you can let AI agents run without a human watching every step — because the failure modes are legible and cheap. You reserve human attention for the small set of actions that genuinely cannot be undone.
A concrete scenario: the remediation AI agent
Consider an agentic AI that keeps a server fleet healthy. It reads host telemetry, spots a drifted configuration or a missing patch, and fixes it.
The naive design gives the AI agent change authority on production hosts and hopes the model is right. One poisoned log line or ticket, or one confident misreading of a host's state, and a production firewall rule or service configuration changes underneath you. Prompt injection through content an agent processes is the top risk in the OWASP Top 10 for LLM Applications, and an infrastructure agent reads untrusted content all day.
The reversible-by-design version classifies each step:
- Reading telemetry and scoring host posture is trivially reversible. The AI agent runs unattended.
- Opening a finding or a ticket is reversible with cost. A human can dismiss it.
- Changing a host (a firewall rule, a service configuration, a patch and reboot) is where enforcement bites. Some changes have a clean undo; a reboot mid-transaction or a deleted volume does not.
Host changes do not execute on the AI agent's judgment. The pattern is closed-loop remediation, human-approved: the system detects the issue, recommends the fix with its evidence, a human approves it, the change is applied with rollback available, and the host is re-scored so the fix is proven rather than assumed. The AI agent still does most of the work. The part that could take down production is bounded.
Reversal on infrastructure also depends on something permission never captures: the host's state before and after the change. You cannot walk back a change if you do not know what the configuration was, and you cannot call it fixed if you do not re-check the host's posture afterward. That before-and-after record is infrastructure trust applied to remediation.
This is the pattern in right at design time, wrong by Tuesday: a system that was safe when it shipped drifts into danger as the model, the data, and the attackers change. Design-time classification plus runtime enforcement is what survives that drift.
Putting it into practice
You do not need to re-architect everything. Start where the consequence is highest and the volume is manageable.
Inventory actions, not AI agents. List what your AI agents can actually do — every write, every external call, every state change. Most teams are surprised by the length of this list. You cannot bound what you have not enumerated.
Classify by reversibility, not by department. Sort each action into trivially reversible, reversible-with-cost, and irreversible. Be honest about the third bucket. "We could probably fix it manually" means irreversible.
Enforce at the boundary. For irreversible actions, put the control at the point of execution — not in a policy document, not in a training. The enforcement has to be in the path the action takes. This is the argument in the trust assertion: a claim of safety is only worth what it is enforced against.
Measure recovery, not just accuracy. Track mean time to reverse a wrong action alongside model accuracy. Google engineers found that in real-world ML systems only a small fraction is the model code itself; the surrounding infrastructure is where the cost and the failures accumulate. The same holds for AI agents. An accurate AI agent with no undo path is more dangerous than a mediocre one you can always walk back.
Expand only once each stage earns it. Small, reversible steps beat one big bet, because each step teaches you something and none of them can sink you.
Common pitfalls
The usual failure is adopting the label without the discipline. Teams announce "reversible AI" and change nothing about how irreversible actions execute.
Confusing logging with reversibility. An audit log tells you what happened. It does not undo it. Observability and reversibility are complementary, not the same — and confusing them leaves you with a perfect record of an unrecoverable mistake.
Classifying once and forgetting. An action's reversibility changes as your systems change. A delete that used to hit a soft-delete table now hits a hard delete after a migration. Reclassification has to be continuous, which is why we treat compliance evidence as a standing capability, not a fire drill.
Gating on permission alone. Zero-trust identity answers who. It does not answer whether an action should run given its consequence. The two controls stack; neither replaces the other.
Optimizing for demo-friendly autonomy. An AI agent that does everything unattended demos beautifully and fails in production. Stay honest: tie every autonomous action to a reversal path you have actually tested, and retire any autonomy that has not earned it.
The takeaway
Reversibility is the enterprise AI safety control that scales with agentic autonomy instead of fighting it.
- Permission is not enough. Authorizing an action says nothing about whether you can undo it.
- Classify every AI agent action by consequence — trivially reversible, reversible with cost, or irreversible — at design time.
- Enforce at runtime, at the point of execution. Irreversible actions get a human or a hard policy in the path; everything else runs free.
- Prove the recovery on the host. Record state before a change and re-check host posture after it, so a reversal is verified, not assumed.
- Measure mean time to reverse, not just model accuracy. Recovery speed is the real safety metric.
- Reversibility is a speed feature. When most actions are cheap to undo, you can let AI agents run — and reserve scarce human attention for the few actions that genuinely cannot be walked back.
The enterprises deploying agentic AI safely are not the ones with the smartest models. They are the ones who designed for being wrong.
Next: scope what each AI agent may do in the first place — read Least Privilege for AI Agents.
Written by
SAUTERA
Author of the Infrastructure Trust Architecture (ITA) and the Infrastructure Trust Conveyance Mechanism (ITCM) — the standard organizations use to decide whether infrastructure can be trusted.
Follow the work
Read the next one
New perspectives on infrastructure trust and updates to the ITA / ITCM framework, by email.
Occasional. No spam. Unsubscribe anytime.