- Published on
Approval Boundaries for Expensive, External, and Irreversible AI Actions
- Authors

- Name
- Mehdi Akiki
Reference
“Ask a human before dangerous actions” is good advice, but it does not define a safe system.
What is dangerous? What exactly did the person approve? Can the model change the arguments afterwards? Does one approval cover one email or 10,000 emails?
I design approval as an authorization artifact, not as a confirmation modal added at the end.
Approval and authentication answer different questions
Authentication says who the user is. Ordinary authorization says which resources that user may access. Approval says that this person accepted one concrete action under known conditions.
The checks compose:
authenticated actor
+ authorized target
+ approved concrete effect
+ current preconditions
= action may execute
An approval cannot make an unauthorized action valid. If the user cannot normally delete another tenant's project, clicking “Approve” in an agent interface must not grant it.
I classify the effect, not the tool name
One tool may contain low-risk and high-risk operations. A database tool that reads a public catalogue differs from the same connector deleting customer rows.
I score the proposed effect along these dimensions:
| Dimension | Lower risk | Higher risk |
|---|---|---|
| reach | internal read | external communication or publication |
| reversibility | preview or temporary change | deletion, payment, legal submission |
| scale | one bounded object | bulk or open-ended operation |
| cost | fixed small budget | variable compute or financial spend |
| data sensitivity | public metadata | secrets, health, identity, customer data |
| ambiguity | exact target and result | inferred recipient or uncertain scope |
The model's confidence is not a replacement for this classification. A confident model can still propose an irreversible action against the wrong account.
Four practical execution classes
I normally use four classes:
Class 0: automatic read
Class 1: automatic bounded and reversible write
Class 2: explicit approval for a concrete effect
Class 3: prohibited through the agent
Examples depend on the product, but one policy can look like:
| Proposed action | Class | Control |
|---|---|---|
| search permitted internal docs | 0 | authorization and audit |
| create an unpublished draft | 1 | quotas and easy rollback |
| send an external email | 2 | show recipients and body |
| delete a production database | 2 or 3 | break-glass workflow or no agent access |
| export all customer records | 3 | dedicated governed process |
This table belongs in code and policy tests. It should not exist only in a prompt that the model may ignore.
Approval comes after argument resolution
A weak flow asks:
“Allow the agent to clean up old resources?”
The phrase hides the target set, deletion mode, cost, and consequences. A useful flow resolves the arguments first:
Delete 7 preview environments:
env_17, env_22, env_31, env_44, env_52, env_61, env_80
Expected effect: stop containers and discard ephemeral files
Maximum objects: 7
Approval expires: 15 minutes
After approval, the executor must use those same arguments. The model does not get another opportunity to replace 7 with all.
Bind consent to a canonical request
I represent approval roughly as:
type Approval = {
approvalId: string;
actorId: string;
tenantId: string;
action: "environment.destroy";
canonicalArgumentsHash: string;
targetIds: string[];
maximumCount: number;
maximumCostCents?: number;
requiredStateVersion?: string;
expiresAt: string;
consumedAt?: string;
};
The arguments are normalised before hashing: sorted keys, explicit defaults, stable identifiers, and no hidden “current selection.”
At execution I verify:
same actor and tenant
same action type
same canonical arguments
targets still authorized
count and cost inside limits
state precondition still current
approval not expired or already consumed
This blocks approval swapping, late argument mutation, and reuse outside the original scope.
State can change between preview and execution
Approval introduces time. During that time, a draft may become public, an account owner may change, or a resource may receive production traffic.
I bind important approvals to a state version or precondition:
approved: delete environment env_17 while version = 84 and traffic = 0
execute: require version = 84 and traffic = 0
If the check fails, I produce a new preview and request a new decision. Silently applying old consent to new state is a time-of-check/time-of-use bug.
Retries should not ask twice or act twice
The approval ID and the action idempotency key have related but different purposes:
approval ID: evidence of consent
action key: identity of the intended external effect
attempt ID: one network or worker attempt
If the executor times out after the external system accepted the request, a retry uses the same action key and reconciles the result. It must not invent a second effect merely because the approval remains valid.
For this wider retry problem, see Make Retried Agent Actions Idempotent Before Adding Autonomy.
Bulk approval needs a hard envelope
A checkbox saying “approve all” is an invitation to scope drift. I show a summary and bind limits:
targets: 142 exact IDs
action: archive, not delete
external messages: none
maximum estimated cost: €4.20
failure mode: stop after 3 errors
expiry: 10 minutes
If the plan later discovers another 18 targets, those targets are not approved. The system can ask again with a new plan.
For very large effects, I add staged execution: approve a sample, verify the outcome, then approve a bounded next batch.
Human review must be readable
Raw JSON is exact but often hides meaning. Friendly prose can hide parameters. I present both a human view and a canonical machine record.
The human view answers:
- what will change;
- where and for whom;
- what leaves the system;
- what it may cost;
- whether it can be reversed;
- what happens on partial failure.
Approval fatigue is also a safety defect. If every harmless read opens a modal, people learn to click without reading. Risk classification must remove low-risk noise so high-risk prompts keep their meaning.
Tests I require
I test the boundary without relying on model behaviour:
- Change one target after approval; execution must fail.
- Reuse an approval for another tenant; it must fail.
- Let the approval expire; it must fail closed.
- Change the resource version; it must require a new preview.
- Retry after an ambiguous timeout; only one external effect may exist.
- Ask for 101 actions with a maximum of 100; the executor must reject it.
- Revoke the actor's permission after approval; current authorization must win.
The important boundary is not “AI versus human.” It is proposal versus authority. The model can prepare a useful action, but the executor should receive only the narrow, current, testable authority that the user actually granted.