Human-on-the-Loop Oversight
The Rubber-Stamp Problem: Why Human-in-the-Loop Is a False Promise, and What Should Replace It
Inserting a human at every AI decision point does not create oversight. It creates the illusion of oversight.
By Christopher Hughes
Personal views of Christopher Hughes, independent of any employer, customer, or third party.
"Every leadership conversation I have right now starts with the same question: how is our AI being supervised? In most organisations, the honest answer is that it isn't. What they've built is control theatre."
The moment an organisation deploys an AI agent, a focused tool built to handle a specific workflow or decision, someone in leadership asks the same question: "But who approves it?" The answer, almost universally, is a human. A checkpoint. A gate. Human-in-the-Loop, or HITL: the principle that a person reviews and approves AI outputs before anything consequential happens.
It sounds responsible. It is, in most implementations, a fiction.
Not because humans are unreliable, but because the model itself is broken. HITL as commonly deployed treats human attention as an infinite resource that can be rationed out across every AI output, at any hour, under any cognitive load. It cannot. And when the volume of approvals exceeds what genuine scrutiny allows, the human in the loop stops being an oversight mechanism. They become a rubber stamp with a pulse. The compliance checkbox gets ticked. The safety signal reaches leadership. The risk remains.
The Problem Hiding in Plain Sight
Automation bias, the tendency to over-trust automated systems and accept their outputs without sufficient scrutiny, is not new. Psychologists Raja Parasuraman and Dietrich Manzey documented it in a foundational 2010 paper: when humans work alongside automated systems, they systematically over-trust the machine's outputs, particularly under cognitive load or when multiple tasks compete for attention.[1] They called it complacency. The automation handles the routine, so the human disengages, not consciously, not maliciously, just naturally. Attention flows to where it is needed. When the machine rarely fails, the human stops checking whether it has.
What is new is the scale at which this plays out inside organisations that have deployed AI agents into their workflows. The approvals are not coming once a day. They are coming in dozens, or hundreds. Each one looks the same. Each one takes 30 seconds. Each one carries the implicit message: the AI has already done the work; your job is to confirm it.
The research on software development makes the dynamic visible. Studies of developers using AI coding assistants such as GitHub Copilot consistently show acceptance rates of suggested code ranging from roughly 20 to 30% of all suggestions shown.[2][3][4][5] That might sound like developers are exercising judgement: accepting some, rejecting others. It is not. An acceptance rate is not a measure of scrutiny. It is a measure of throughput.
Healthcare figured this out the hard way. In intensive care units, physiological monitors generate an average of 152 alarms per bed per day.[6] Between 80 and 99% of those alarms are false positives, triggered by patient movement, loose sensors, or equipment artefacts rather than genuine clinical events.[7] Nurses, trained to respond to every alarm, facing hundreds of false signals per shift, began doing what any rational person would do: silencing alarms, lowering thresholds, and in some documented cases, disabling monitoring equipment altogether.[7] The alarm fatigue literature does not describe negligent nurses. It describes a system designed without any account of human attention limits, producing the opposite of its intended effect.
This is what HITL looks like when it breaks.
Not All Decisions Are the Same
The deeper problem with HITL as a blanket policy is that it applies identical oversight to decisions that have nothing in common.
Routing a customer support ticket to the right team does not carry the same consequence as adjusting a pricing rule that affects 100,000 contracts. Flagging a candidate's CV for an initial screen does not require the same scrutiny as authorising a vendor payment above a material threshold. These decisions differ on two axes that matter: consequence severity and reversibility.
A misrouted support ticket costs a delay. It can be corrected in minutes. An incorrect pricing change in a high-volume environment can take weeks to unwind and may require customer remediation, regulatory disclosure, and finance reconciliation. The risk profile is categorically different. Treating them identically (one human, one approval, one click) is not risk management. It is risk theatre.
The right question is not "should a human approve this?" It is "what is the cost of this decision being wrong, and can it be undone quickly?" Those two factors determine the appropriate level of oversight. A consequential, irreversible decision warrants intensive human scrutiny. A low-stakes, easily reversible one warrants an audit trail and a circuit breaker, not a human gateway.
What Other Industries Already Know
Aviation did not wait for AI to solve this problem. The same dynamic, skilled humans over-trusting automated systems they barely needed to engage with, nearly destroyed commercial aviation's safety record in the 1990s.
The crash of Air France Flight 447 in June 2009 is the case study the industry returns to most often. The aircraft lost its airspeed indicators over the Atlantic. The autopilot disconnected. Three experienced pilots, handed manual control of a perfectly flyable aircraft, failed to recognise they were in a stall and flew the plane into the ocean. The official investigation by the French Bureau d'Enquêtes et d'Analyses found that none of the crew had been trained to fly the aircraft manually at high altitude. The automation had done it for so long, and so reliably, that the skill and the situational awareness had quietly drained away.[8]
Aviation's response was not to remove automation. It was to redesign the relationship between humans and automation. Crew Resource Management training was updated to make automation management an explicit competency, not an assumed one. Pilots are now trained specifically to recognise automation complacency, to fly manually on a rotation basis to maintain proficiency, and to understand the boundaries at which the system's confidence should not be trusted. The human is not in the loop at every moment. The human is on the loop: monitoring, calibrated, ready to intervene where their judgement genuinely adds something the automation cannot provide.[9]
Financial markets learned the same lesson differently. Exchanges do not require a human to approve every trade before it executes. At the volume modern markets operate, that would be physically impossible. Instead, they run automated rules continuously, and only halt or escalate when those rules are breached. The SEC's market-wide circuit breakers trigger automatic trading halts when the S&P 500 falls 7%, 13%, or 20%, thresholds set by regulators as indicators that something genuinely exceptional is occurring.[10] Below those thresholds, the market runs. Above them, humans engage. The oversight is calibrated to the signal, not inserted at every transaction.
Neither aviation nor financial markets achieved safety by putting a human in front of every automated output. They achieved it by designing intelligent thresholds that determine when human judgement is actually required. That is the model operations leaders need to borrow.
Human-on-the-Loop: What Genuine Oversight Looks Like
The alternative to Human-in-the-Loop is not no oversight. It is better-designed oversight: Human-on-the-Loop, or HOTL.
The distinction is architectural. HITL puts humans in the approval pathway of every output. HOTL puts humans in a monitoring and exception role, engaged actively when the AI system itself flags uncertainty or when a decision crosses a consequence threshold. The human's time and attention are directed to where they demonstrably change the outcome.
Four design principles make HOTL work in practice.
Confidence-gated escalation
AI agents can be designed to report their own uncertainty, routing low-confidence outputs to human review while processing high-confidence ones autonomously. This is not a new concept; spam filters have done it for decades. The standard of proof for human review is calibrated to the stakes, not applied uniformly.
Consequence-weighted routing
Not all decisions escalate on the same trigger. A decision threshold for auto-approval on a customer communications task is different from the threshold on a financial commitment or a personnel action. These thresholds are set by humans, reviewed regularly, and adjusted as the system's track record warrants.
A genuine audit trail
The purpose of an audit trail is not to create a paper record that a human approved something. It is to make oversight verifiable and actionable. An audit trail that shows what the AI decided, on what inputs, with what confidence, and what a human did with that, including whether the human modified, rejected, or overrode the output, gives the organisation the data to improve both the AI and the oversight model over time.
Friction as a feature
Most approval interfaces are designed to be frictionless: one click, one screen, minimal cognitive load. But frictionless approval is the mechanism of the rubber stamp. When a decision is consequential, the interface should make approval genuinely difficult, requiring the reviewer to confirm they have read the relevant inputs, to categorise the nature of their approval, and to record the basis for the decision. Friction is not a UX failure. It is a governance mechanism. The goal is not to make humans faster at approving. It is to make humans actually approve.
The Incentive Problem Nobody Wants to Name
Implementing HOTL is technically straightforward. The harder problem is organisational.
HITL survives not because it works, but because it satisfies incentives that have nothing to do with risk management. It satisfies legal and compliance functions that want documented human approval. It satisfies boards and regulators who want to see that "a human checks the AI." It satisfies executives who want to be able to say, when something goes wrong, that a human was responsible. HITL is, above all, a liability distribution mechanism dressed up as a safety mechanism.
That is not a reason to keep it. It is a reason to be honest about what it is doing, and to design something that actually works in its place.
The transition to HOTL requires organisations to be explicit about something HITL lets them avoid: a clear, documented risk tolerance for each category of AI decision. That tolerance has to be set by a human, signed off by leadership, and reviewed as the AI system's track record evolves. It cannot be set once and forgotten. The oversight model for an AI agent that has processed 10,000 decisions with a documented error rate of 0.3% should be different from the oversight model for one deployed 6 weeks ago. Gartner's May 2026 research puts a number on what happens when organisations skip this step: 40% of enterprises will demote or decommission AI agents by 2027 due to governance gaps that only became visible after production incidents.[11] Uniform governance applied indiscriminately is its own failure mode.
Approval fatigue does not announce itself. It just quietly turns a governance mechanism into a keystroke.
The Next Step
If your organisation has deployed AI agents into any operational workflow, run this test. Pull the last month of approvals from your HITL process. Ask how long each review took, how often approvals were modified or rejected, and whether the humans reviewing those outputs had access to the inputs the AI used. If average review time is under 90 seconds, rejection rates are below 5%, and the reviewers cannot tell you what inputs drove the AI's recommendation, you do not have oversight. You have the performance of oversight.
The fix is not to add more checkpoints. It is to redesign which decisions need human judgement at all, build thresholds that route only those decisions to human review, and design the review interface to make genuine scrutiny possible rather than optional.
The human is not the last line of defence. The human is the most expensive cognitive resource in the system. Deploy them accordingly.
Sources
- [1] Raja Parasuraman and Dietrich Manzey, "Complacency and Bias in Human Use of Automation: An Attentional Integration," Human Factors, Vol. 52, No. 3, 2010, pp. 381–410. DOI: 10.1177/0018720810376055.
- [2] Christian Bird et al., "Taking Flight with Copilot," Communications of the ACM, Vol. 66, No. 6, May 2023. DOI: 10.1145/3589996.
- [3] "Measuring GitHub Copilot's Impact on Productivity," Communications of the ACM, Vol. 67, No. 3, March 2024. DOI: 10.1145/3626246.
- [4] Thomas Dohmke, Marco Iansiti, and Greg Richards, "Sea Change in Software Development: Economic and Productivity Analysis of the AI-Powered Developer Lifecycle," arXiv:2306.15033, June 2023. Large-scale dataset: n = 934,533 users; modal acceptance rate approximately 27–30%.
- [5] Gal Bakal et al., "Experience with GitHub Copilot for Developer Productivity at Zoominfo," arXiv:2501.13282, January 2025. Enterprise replication; acceptance rates range from approximately 21% to 34% depending on developer experience level.
- [6] Maria Cvach, "Monitor Alarm Fatigue: An Integrative Review," Biomedical Instrumentation & Technology, Vol. 46, No. 4, 2012, pp. 268–277. DOI: 10.2345/0899-8205-46.4.268.
- [7] "Exploring ICU nurses' response to alarm management and strategies for alleviating alarm fatigue: a meta-synthesis and systematic review," PMC, 2025. False alarm rates of 80–99%; documented unsafe practices including silencing and disabling alarms.
- [8] Bureau d'Enquêtes et d'Analyses (BEA), "Final Report on the Accident on 1st June 2009 to the Airbus A330-203 registered F-GZCP operated by Air France flight AF 447 Rio de Janeiro, Paris," July 2012.
- [9] Robert L. Helmreich, Ashleigh C. Merritt, and John A. Wilhelm, "The Evolution of Crew Resource Management Training in Commercial Aviation," FAA Human Factors Research and Engineering Division, archived November 2022.
- [10] U.S. Securities and Exchange Commission, "SEC Approves Proposals to Address Extraordinary Volatility in Individual Stocks and Broader Stock Market," Press Release, June 2012. Circuit breaker thresholds of 7%, 13%, and 20% confirmed via NYSE, "Report of the Market-Wide Circuit Breaker Working Group."
- [11] Gartner, "Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure," Press Release, May 26, 2026.
Put this paper to work
Discuss what this changes for your organisation