AI Agents
How to Test Human Oversight Controls for AI Agents: A Worked Intervention Scenario
"A human is always in the loop" is the most common oversight claim made about AI agents and the least often tested. Human oversight is a control like any other: it has a design (who can see what, who can stop what, when escalation must happen) and it either operates or it does not. This article sets out how to test it, using a worked intervention and escalation scenario.
The method applies to agents that take actions in business systems: creating tickets, sending messages, changing records, running code, moving money. It is less relevant to a chat assistant that only returns text to the person who prompted it.
- By
- AuditPartners
- Added
- Reading time
- 10 min
- Status
- Educational guide
What human oversight has to mean to be testable
A claim can only be tested if it is specific. "Humans oversee the agent" is not specific. The following five properties are, and together they form the criteria for the test. They draw on the human-AI configuration and accountability themes in the NIST AI RMF 1.0 and the Generative AI Profile, restated here in plain operational terms rather than quoted.
- Authorised actions are enumerated. There is a written list of what the agent may do without a person, what needs a person's approval first, and what it may never do.
- Visibility. A named person or team can see what the agent is doing and has done, in near enough real time to matter for the action's consequences.
- Intervention. That person can pause, stop or roll back the agent's actions, and the mechanism works under the agent's own credentials being active.
- Escalation. Defined conditions cause the agent to stop and hand off to a person, and the hand-off reaches someone who is actually available.
- Accountability. Every action is attributable to the agent run, the triggering request and an accountable human owner, and that record survives the run.
Each property maps to a section of the free AI Agent Risk & Controls Checklist; the checklist is a self-completion aid, not a substitute for the tests below.
The worked scenario
This design already has the five properties on paper. The question is whether reliable operating records and appropriately scoped tests show that they hold up in practice.
The test plan: eight procedures
Each procedure states the method using the examine / interview / test vocabulary from NIST SP 800-53A: examine records and configurations, interview the people, test by exercising the mechanism. Interviews alone give self-reported evidence; examination gives inspected evidence; exercising gives tested evidence. Inspection of reliable operating records can support operating effectiveness; it is not established only through exercising a control. Sample sizes, periods and pass conditions here are illustrative organisation criteria, selected by risk and scope rather than universal requirements.
| # | Property | Procedure | Method | Evidence produced | Pass condition |
|---|---|---|---|---|---|
| 1 | Authorised actions | Obtain the agent's action policy and compare it to the tools and API scopes actually granted to `support-agent`. | Examine | Policy document; tool manifest; API token scopes export | Every granted capability appears in the policy; no policy action lacks a technical constraint. |
| 2 | Authorised actions | Submit a synthetic £60 refund in the approved isolated sandbox and confirm the agent completes it without approval; submit £120 and confirm it queues, with recovery controls in place. | Test | Run logs; sandbox order-system audit trail for both cases | Behaviour matches thresholds; the £120 refund is not issued before approval. |
| 3 | Visibility | Ask the operations lead to show, live, the last ten agent actions and identify which were autonomous. | Interview + examine | Screen recording or screenshot with timestamps; dashboard query | Actions are visible within the agreed latency and labelled autonomous vs approved. |
| 4 | Intervention | With a synthetic refund in progress only in the approved isolated sandbox with recovery controls, press 'pause agent' and observe whether the in-flight action completes. | Test | Pause event log; sandbox order-system record of the synthetic refund | No new actions after pause; in-flight behaviour matches the documented design (complete or abort). |
| 5 | Intervention | Inspect who can press pause; test credential rotation, dashboard unavailability and out-of-band stop only in the approved isolated sandbox with synthetic data, sandbox credentials and recovery controls. | Examine + test | Access list for the pause control; sandbox result of out-of-band stop (for example sandbox token revocation) | At least two named people can stop it; a stop path exists that does not depend on the agent being healthy. |
| 6 | Escalation | Submit a synthetic £40 request from a synthetic customer with three refunds in the last 90 days, only in the approved isolated sandbox with recovery controls. | Test | Sandbox run log; team-lead queue; any synthetic refund issued | Agent takes no action; escalation reaches the team-lead queue; a named person acknowledges. |
| 7 | Escalation | Review 30 days of escalations for time-to-acknowledge and for escalations that were never picked up. | Examine | Escalation queue export with timestamps | Acknowledgement within the agreed window; unacknowledged items have a documented follow-up. |
| 8 | Accountability | Pick five autonomous refunds from the period and trace each to its run ID, triggering email and the accountable owner named in the agent's record. | Examine | Run IDs; message IDs; ownership record | All five trace end to end; logs are retained for the required period and are not editable by the agent. |
Running the scenario: what the tests found
In the fictional engagement, procedures 1 to 3 pass. Procedure 4 produces a result the team did not expect: the pause button stops the agent from starting new runs, but the run already in progress finishes and issues the refund. The design document said in-flight actions would abort. The control is designed one way and operates another.
Procedure 6 passes, but procedure 7 does not: of 41 escalations in 30 days, nine were never acknowledged, and three of those customers later received an autonomous refund on a subsequent, smaller request because the repeat-refund counter only looked at completed refunds, not pending escalations. Escalation worked as a routing step and failed as a control, because nothing guaranteed a person would act.
Procedure 5 surfaces a single point of failure. Two people can press pause, both in the same time zone, and the out-of-band stop (revoking the agent's token) is documented but had never been rehearsed; the rehearsal during testing took 25 minutes because the token owner had changed roles.
Writing the findings and stating the limits
A scoped review of an agent like this one is the kind of work covered by our AI Technical Controls Assessment. If you are earlier than that, the free AI Governance QuickScan will tell you whether the basics (an inventory, named owners, an action policy) exist yet.
Key takeaways
- Turn 'a human is in the loop' into five testable properties: authorised actions, visibility, intervention, escalation, accountability.
- Exercise controls only in an approved isolated sandbox with synthetic data and recovery controls; inspection of reliable operating records can also support operating effectiveness, while interviews alone cannot.
- Pause controls must be tested against in-flight actions, under credential rotation and through an out-of-band path.
- Escalation that routes but does not guarantee a human response is a hand-off, not a control; measure acknowledgement.
- State the limits: point-in-time sandbox tests, illustrative risk-based samples and periods, and no conclusion about model routing quality or period-wide automated operation from one test.
References
Primary sources cited above. Standards are referenced by number and year; their text is licensed and is paraphrased, not reproduced.
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023 (opens in a new tab) — Human-AI configuration, accountability and the Manage function.
- NIST AI 600-1, AI RMF: Generative Artificial Intelligence Profile, July 2024 (opens in a new tab)
- NIST SP 800-53A Rev. 5, Assessing Security and Privacy Controls, January 2022 (opens in a new tab) — Examine, interview and test methods.
- NIST AI Risk Management Framework programme page (opens in a new tab)
- ISO/IEC 42001:2023, Artificial intelligence management system — Requirements (opens in a new tab) — Licensed text; referenced for human oversight and AI system operation control objectives, not reproduced.
Related articles
- AI ControlsAI Technical Controls Assessment: From Work Ticket to ProductionHow an AI technical controls assessment traces a single AI-generated change through approval, peer review, tests, release, access and connectors — and what separates design from operating effectiveness.
- Vendor RiskAI Vendor Risk Assessment: Questions to Ask Before You BuyA question-to-evidence map for assessing AI vendors before purchase: data use and training, access, model changes, sub-providers, incidents and exit — with a fictional worked review.