GuideAI Agents
Human-in-the-loop design: deciding what an AI agent may do without asking
Human-in-the-loop is not a single setting. It is a decision made per action: which actions wait for approval, which are checked afterwards and which surface only when something looks wrong. This guide gives a risk-tiering method based on reversibility and impact, the evidence an approval screen must show, what to do when nobody answers, and the measures that justify moving an action to lighter oversight.
On this page
- Oversight modes and what each one costs
- Risk tiers for the actions an agent can take
- What the approval screen has to show
- When nobody answers: timeouts and escalation
- Turning reviewer decisions into evaluation data
- Widening autonomy one action type at a time
- Rubber-stamping and other ways oversight fails
- Where law and standards expect human oversight
- A hypothetical utility service agent, tiered action by action
- Questions and answers
- Sources
Oversight modes and what each one costs
Each mode trades reviewer time against the chance that a mistake reaches a customer or a system of record.
- Approve-first
- The agent prepares the action and waits; nothing happens until a person approves, edits or rejects it. Safest and slowest, and the default for new or high-impact actions.
- Review-after
- The agent acts within limits, and a person reviews a sample or every action afterwards, with a tested way to reverse it. Suited to reversible actions with a good record.
- Exception-only
- The agent acts alone and escalates only when a rule, a confidence threshold or a limit is triggered. Reserved for low-impact actions backed by strong evidence.
- Stop authority
- The power of a named person to pause the agent, or one type of action, immediately, whichever mode is in force. Every mode needs it.
Risk tiers for the actions an agent can take
Classify each tool and action, not the agent as a whole. When in doubt, place an action in the higher tier.
| Tier | Typical actions | Starting oversight | What the reviewer must see |
|---|---|---|---|
| Read | Look up a record, search documents, check a status | None beyond access controls and logging | Nothing per action; a periodic review of what was accessed |
| Draft | Prepare a reply, a form or a proposed record change that is not yet sent or saved | Review-after or exception-only, since nothing has left the system | The draft beside the sources it used |
| Reversible write | Update an internal field, assign a ticket, add a note | Approve-first at launch, moving to review-after on evidence | A before-and-after view of the record and the reason for the change |
| Irreversible or external | Send a message to a customer, delete data, submit to a third party | Approve-first | The exact content to be sent or removed, the recipient and the evidence |
| Financial, legal or customer-impacting | Issue a refund, change a contract term, alter eligibility or credit | Approve-first, often with a second approver above a set limit | The policy basis, amounts, the affected party and any exceptions the agent noticed |
What the approval screen has to show
A reviewer can only catch what the screen makes visible, so design the approval moment as carefully as the agent.
When nobody answers: timeouts and escalation
An approval queue without timeout rules becomes a backlog that people clear in bulk, which defeats its purpose.
- If
A reversible, low-impact action waits past its response target
ThenEscalate to a backup reviewer or a team queue, and tell the requester the item is pending.
A short delay is better than an unreviewed action.
- If
An irreversible, financial or customer-facing action times out
ThenNever approve by default. Escalate, and if the deadline passes, stop safely and tell the requester that a person will follow up.
- If
The action carries an external deadline, such as a contractual or regulatory response time
ThenSet the timeout well inside that deadline and route to a reviewer group with cover across shifts.
- If
Items time out regularly
ThenTreat it as a design fault: the tiering may be too strict, the volume too high or the screen too slow to review.
Chronic timeouts push reviewers toward bulk approval, which removes the oversight you designed.
Turning reviewer decisions into evaluation data
Every approval, edit and rejection is a labeled example of what the agent should have done. Capture the decision, a reason code from a short fixed list and the final version of any edited content. Within a few weeks this becomes the most realistic test data the team has.
Use it in two ways. Edited outputs become reference answers in the agent's evaluation set, and rejection reasons point to the categories that need new cases or a rule. Our guide to building an LLM evaluation set explains how to keep them as regression cases.
While an action waits, the agent should hold the task state and avoid any dependent step that assumes approval.
Widening autonomy one action type at a time
Agree promotion criteria before launch
For each action type, write down the evidence that would justify lighter oversight: a minimum volume of reviewed actions, an acceptable edit rate, no high-severity errors and an acceptable time to decision.
Run approve-first and record everything
Start new and higher-tier actions in approve-first mode, logging approvals, edits, rejections with reasons and how long each decision took.
Grade the errors you catch
Classify each edit or rejection as cosmetic, needing correction or potentially harmful. A low edit rate that includes one harmful error is not a case for promotion.
Promote to review-after with sampling
When the criteria are met, let the action execute and review a sample afterwards, starting with a high sampling rate and a tested way to reverse it.
Demote on evidence
A harmful error, a change of model or prompt, or a new kind of input sends the action back to approve-first until fresh reviews support promotion again.
Rubber-stamping and other ways oversight fails
Automation bias
Early signalApproval times fall to seconds and edits disappear, while errors still surface elsewhere.
MitigationSeed known-flawed items into the queue, rotate reviewers and track whether the seeded errors are caught. The EU AI Act expects people overseeing high-risk systems to stay aware of this tendency to over-rely on system output1.
Reviewer overload
Early signalQueues grow and items are approved in bulk at the end of a shift.
MitigationSet workload limits per reviewer, move low-tier actions out of the queue and staff to the actual volume.
Screens that hide the evidence
Early signalReviewers approve the agent's summary without opening the record.
MitigationShow diffs and sources inline, so that checking is quicker than skipping.
Unclear accountability
Early signalNobody can say who approved a harmful action, or under which policy.
MitigationRecord the reviewer, time, evidence shown and policy version for every decision, and make the process owner, not the agent team, responsible for the tiering.
Where law and standards expect human oversight
A hypothetical utility service agent, tiered action by action
Questions and answers
Does requiring approval defeat the purpose of an AI agent?
No. Even in approve-first mode the agent does the gathering, checking and drafting, so the reviewer makes a decision in moments instead of doing the whole task. Approvals also produce the evidence needed to move low-risk actions to lighter oversight. The aim is to spend human attention where errors would matter, not to remove it everywhere.
Who should approve the actions an AI agent proposes?
The people who would make the decision if there were no agent: the process owner's team, with authority matching the action. Financial or legal actions above a limit may need a second approver. Avoid making the engineers who built the agent its approvers, because they can judge technical correctness but not the business decision.
How do we audit decisions made with an agent's help?
Log, for every action, the input, the evidence retrieved, the proposed action, its tier, the reviewer's decision and reason, the change actually executed and the system's response, with the policy and prompt versions. An auditor should be able to reconstruct why an action happened and who agreed to it without asking the team.
Can a second AI model approve an agent's actions instead of a person?
A second model can screen proposals, flag likely errors and reduce what reaches people, which is useful in the lower tiers. It should not be the final approver for irreversible, financial, legal or customer-impacting actions, because both models can share the same blind spots and neither carries accountability. Treat it as a filter in front of human review.
Sources
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Articles 14 and 26 — EUR-Lex · checked 10 October 2026
- Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22 — EUR-Lex · checked 10 October 2026
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Govern function, subcategory 3.2) — National Institute of Standards and Technology · checked 10 October 2026