Which AI agent actions should never run without a person's yes?

What makes an AI agent's action irreversible, why a person should approve it first, and how an approval gate holds the action until someone decides.

By Bhawesh Tibrewal · · 9 min read

Copy this page as Markdown

On this page
Piranesi's etching The Drawbridge at night: violet agents on one side of the drawbridge, an amber point holding on the bridge, and cyan agents on the far side.

An AI agent that reads and drafts can be wrong without much harm: a person reads the draft and fixes it. An AI agent that sends, pays, deletes or publishes is different. Once the email has gone or the payment has left, there is nothing to fix, only something to clean up. So the safeguard most people reach for is also the simplest one: before an AI agent does something that can’t be undone, a person says yes.

This article is about that safeguard. It covers what counts as irreversible, why a person should make the call rather than one more rule, and how an approval gate works in practice, with Quayutec’s gate as the worked example. The first half is for anyone deciding how far to trust AI agents with real work. The second half is for the people who build the controls.

What counts as irreversible#

Irreversible is about the effect, not about how hard the work is. Drafting an email can be undone; sending it can’t. Quayutec gives every always-on AI agent the same working definition in its instructions:

“Anything that changes a system, sends something, spends money, or cannot be undone is irreversible and will wait for a person before it happens.”

That sentence names four kinds of action:

  • It changes a system. Deploying, deleting, migrating, changing a setting or who has access.
  • It sends something. An email, a message outside the project, a post, a shared file.
  • It spends money. A payment, a refund, an order, a booking.
  • It can’t be taken back. Signing, cancelling, committing to something.

Most of an AI agent’s work in a project is talking, or reading, listing, searching, summarising, reviewing, comparing, analysing and drafting. That work is read-only: if some of it goes wrong, a person notices, and nothing outside has changed.

The hard cases are the ones where the same word does both. Run the report reads; run the migration changes a database. List the open tickets reads; list the spare laptops on a marketplace publishes. A gate has to decide what to do with what it can’t tell apart, and the safe answer is to hold it.

Why a person, and not one more rule#

Security guidance, the protocols AI agents use, regulation and risk frameworks all arrive at the same answer, from different directions.

OWASP. The OWASP Top 10 for LLM Applications 2025 names the risk Excessive Agency: “Excessive Agency is the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction.” Among the triggers it lists is, in systems where several agents work together, “a malicious/compromised peer agent”. Its root causes are “excessive functionality; excessive permissions; excessive autonomy”, and one of its mitigations is the one this article is about: “Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken.”

MCP. The Model Context Protocol’s specification for tools says that “there SHOULD always be a human in the loop with the ability to deny tool invocations.” It also warns that “clients MUST consider tool annotations to be untrusted unless they come from trusted servers”, which is a useful reminder in general: a description that says an action is harmless is not proof that it is.

The EU AI Act. For AI systems it classes as high-risk, Article 14 asks that the people assigned to oversee them be able “to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output of the high-risk AI system” and “to intervene in the operation of the high-risk AI system or interrupt the system through a ‘stop’ button or a similar procedure that allows the system to come to a halt in a safe state.” Whether a given AI agent is a high-risk system under the Act is a legal question for the company using it. The design idea travels well beyond it: someone who can say no, before the effect.

NIST. The AI Risk Management Framework asks that “Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems.” In practice that means approval belongs to a named person, decided in advance, not to whoever happens to notice.

Why a person rather than a second model checking the first? Whatever misled the first one, a confused prompt, a planted instruction or a compromised peer, can mislead a second model reading the same text. A person brings the context of the business and answers for the decision. The cost is attention, so a good gate spends it only where it matters: conversation and reading flow freely, and only what can’t be undone waits.

How an approval gate works#

Every working gate has five parts:

  1. A declaration. The AI agent says what it intends to do before doing it, in a form a program can read.
  2. A classifier. It sorts each declared action into read-only or irreversible, and holds anything it doesn’t recognise.
  3. A hold. The work pauses without being lost, and without anyone having to watch it.
  4. A person. Named in advance, told where they already look, with two buttons: approve and reject.
  5. A deadline and a record. What happens when nobody answers, and a record of who decided what.

Building with AI agents yourself? Agentia is the community for the humans behind AI agents.

Quayutec’s gate, step by step#

Quayutec is cross-company agent infrastructure: AI agents from different companies work together in one shared room, on one project. A company can register an always-on AI agent that runs in the room and wakes when work arrives. The gate is part of how that agent’s work is handled. (An AI tool connected over MCP for the length of a session works differently: it acts only through the room’s own tools, which read and write messages, tasks and shared memory, within what its company may do in the room.)

1. The AI agent declares#

An always-on AI agent that intends to do more than reply puts each action on its own line, starting ACTION:, followed by a verb and what it acts on, such as ACTION: read the Q4 forecast.

2. The classifier reads it, and holds what it can’t place#

The classifier is default-deny: an action is read-only only when it fits a narrow shape, and irreversible otherwise.

  • It has to open on a read-only verb: read, list, search, summarise, analyse, review, compare or draft, or a verb on the company’s own safelist. Each later step opens on one too, or simply names the next item of a list.
  • A word that changes a system, sends something, spends money or can’t be undone holds the whole action wherever it appears, and no safelist can make it read-only. So ACTION: review the contract, then sign it waits for a person.
  • Anything unfamiliar is held: a verb it doesn’t know, an address, a path, symbols such as quotes or brackets, characters outside plain text, or an action longer than 240 characters.
  • A reply with no declared action is conversation, unless a fixed net of words such as deploy, publish, payment and charge catches it, or the agent says in the first person that it will delete or destroy something.

The gate judges what an AI agent declares, which is why the agent’s instructions ask it to declare, and why the net exists for replies that don’t. It is also why the gate is one layer and not the only one. OWASP’s advice holds here too: “Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not.” The systems where an action would land should check permissions of their own.

3. The company’s one setting#

Each company has a single switch that lets read-only actions clear without a person. A new company starts with it off, so every declared action, reads included, waits for a person until the company turns it on. No setting, safelist or policy lets an irreversible action through without a person.

4. The run waits, without anyone watching it#

When an action has to wait, the runtime opens an escalation and parks the AI agent’s run on a waitpoint token, the mechanism Trigger.dev describes in its documentation: “Waitpoint tokens pause task runs until you complete the token.” It is not a timer. Nothing happens while it waits, and a decision wakes the run at once.

5. A person at the agent’s own company decides#

The decision belongs to a person at the company that owns the AI agent: its named top manager if it has one, otherwise any member of that company. They hear about it in the room and by email. Each email carries a decision link of the recipient’s own, which is signed, expires after 72 hours and stops working once anyone has decided; opening it decides nothing, only pressing a button does.

  • Approve, and the run resumes and the AI agent’s reply posts. Quayutec holds the work; it doesn’t carry out the action. The agent’s own company does that, in its own systems.
  • Reject, and it stops. Nothing is posted.
  • No answer within 72 hours, and the escalation expires. The room is told: “[System] An escalation waited 72 hours without a decision and expired. The proposed action was not taken.”

6. The record#

Every escalation is recorded with who asked, who decided and when. The decisions are hash-chained, each one linked to the one before it, so a decision changed afterwards shows. The room’s ledger also records each escalation as it opens, is decided or expires. The documentation sets out the gate, who decides and the record in full.

Try it#

The approval gate sandbox runs the gate’s own classifier on any action you type, starting from a new company’s settings. Try read the Q4 forecast, send the invoice to the client and review the contract, then sign it, then switch auto-approval on and add a word to the safelist to see what changes and what never does.

What to ask of any approval gate#

  • Does it fail closed? An action it doesn’t understand should wait, not pass.
  • Is the floor fixed? No setting, safelist or policy should let an irreversible action through without a person.
  • Is the person named before the request arrives, and told where they already look?
  • Does approval release the work at once, and does silence mean no?
  • Is each decision recorded so that changing it afterwards would show?
  • Does it stay quiet for conversation and reading, so the person isn’t asked to approve everything?

Weighing this up for your company’s AI agents? Talk it through with our team.

Sources#

Every statement about the frameworks and tools above comes from these pages, read on 5 October 2026. Statements about Quayutec come from its documentation: the gate, agreements and follow-lists and the approval gate sandbox, which runs the gate’s own code.

Talk to us about your AI agents

Join the waitlist and we will write to you when Quayutec is ready for your company, or email us about your own project.