Mehdi Akiki
Published on

Approval Boundaries for Expensive, External, and Irreversible AI Actions

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Reference

“Ask a human before dangerous actions” is good advice, but it does not define a safe system.

What is dangerous? What exactly did the person approve? Can the model change the arguments afterwards? Does one approval cover one email or 10,000 emails?

I design approval as an authorization artifact, not as a confirmation modal added at the end.

Approval and authentication answer different questions

Authentication says who the user is. Ordinary authorization says which resources that user may access. Approval says that this person accepted one concrete action under known conditions.

The checks compose:

authenticated actor
    + authorized target
    + approved concrete effect
    + current preconditions
    = action may execute

An approval cannot make an unauthorized action valid. If the user cannot normally delete another tenant's project, clicking “Approve” in an agent interface must not grant it.

I classify the effect, not the tool name

One tool may contain low-risk and high-risk operations. A database tool that reads a public catalogue differs from the same connector deleting customer rows.

I score the proposed effect along these dimensions:

DimensionLower riskHigher risk
reachinternal readexternal communication or publication
reversibilitypreview or temporary changedeletion, payment, legal submission
scaleone bounded objectbulk or open-ended operation
costfixed small budgetvariable compute or financial spend
data sensitivitypublic metadatasecrets, health, identity, customer data
ambiguityexact target and resultinferred recipient or uncertain scope

The model's confidence is not a replacement for this classification. A confident model can still propose an irreversible action against the wrong account.

Four practical execution classes

I normally use four classes:

Class 0: automatic read
Class 1: automatic bounded and reversible write
Class 2: explicit approval for a concrete effect
Class 3: prohibited through the agent

Examples depend on the product, but one policy can look like:

Proposed actionClassControl
search permitted internal docs0authorization and audit
create an unpublished draft1quotas and easy rollback
send an external email2show recipients and body
delete a production database2 or 3break-glass workflow or no agent access
export all customer records3dedicated governed process

This table belongs in code and policy tests. It should not exist only in a prompt that the model may ignore.

Approval comes after argument resolution

A weak flow asks:

“Allow the agent to clean up old resources?”

The phrase hides the target set, deletion mode, cost, and consequences. A useful flow resolves the arguments first:

Delete 7 preview environments:
env_17, env_22, env_31, env_44, env_52, env_61, env_80

Expected effect: stop containers and discard ephemeral files
Maximum objects: 7
Approval expires: 15 minutes

After approval, the executor must use those same arguments. The model does not get another opportunity to replace 7 with all.

I represent approval roughly as:

type Approval = {
  approvalId: string;
  actorId: string;
  tenantId: string;
  action: "environment.destroy";
  canonicalArgumentsHash: string;
  targetIds: string[];
  maximumCount: number;
  maximumCostCents?: number;
  requiredStateVersion?: string;
  expiresAt: string;
  consumedAt?: string;
};

The arguments are normalised before hashing: sorted keys, explicit defaults, stable identifiers, and no hidden “current selection.”

At execution I verify:

same actor and tenant
same action type
same canonical arguments
targets still authorized
count and cost inside limits
state precondition still current
approval not expired or already consumed

This blocks approval swapping, late argument mutation, and reuse outside the original scope.

State can change between preview and execution

Approval introduces time. During that time, a draft may become public, an account owner may change, or a resource may receive production traffic.

I bind important approvals to a state version or precondition:

approved: delete environment env_17 while version = 84 and traffic = 0
execute:  require version = 84 and traffic = 0

If the check fails, I produce a new preview and request a new decision. Silently applying old consent to new state is a time-of-check/time-of-use bug.

Retries should not ask twice or act twice

The approval ID and the action idempotency key have related but different purposes:

approval ID:       evidence of consent
action key:        identity of the intended external effect
attempt ID:        one network or worker attempt

If the executor times out after the external system accepted the request, a retry uses the same action key and reconciles the result. It must not invent a second effect merely because the approval remains valid.

For this wider retry problem, see Make Retried Agent Actions Idempotent Before Adding Autonomy.

Bulk approval needs a hard envelope

A checkbox saying “approve all” is an invitation to scope drift. I show a summary and bind limits:

targets: 142 exact IDs
action: archive, not delete
external messages: none
maximum estimated cost: €4.20
failure mode: stop after 3 errors
expiry: 10 minutes

If the plan later discovers another 18 targets, those targets are not approved. The system can ask again with a new plan.

For very large effects, I add staged execution: approve a sample, verify the outcome, then approve a bounded next batch.

Human review must be readable

Raw JSON is exact but often hides meaning. Friendly prose can hide parameters. I present both a human view and a canonical machine record.

The human view answers:

  • what will change;
  • where and for whom;
  • what leaves the system;
  • what it may cost;
  • whether it can be reversed;
  • what happens on partial failure.

Approval fatigue is also a safety defect. If every harmless read opens a modal, people learn to click without reading. Risk classification must remove low-risk noise so high-risk prompts keep their meaning.

Tests I require

I test the boundary without relying on model behaviour:

  1. Change one target after approval; execution must fail.
  2. Reuse an approval for another tenant; it must fail.
  3. Let the approval expire; it must fail closed.
  4. Change the resource version; it must require a new preview.
  5. Retry after an ambiguous timeout; only one external effect may exist.
  6. Ask for 101 actions with a maximum of 100; the executor must reject it.
  7. Revoke the actor's permission after approval; current authorization must win.

The important boundary is not “AI versus human.” It is proposal versus authority. The model can prepare a useful action, but the executor should receive only the narrow, current, testable authority that the user actually granted.