Human-in-the-Loop AI Safeguards Are Quietly Failing, Researchers Warn
A paper from Hugging Face's ethics team argues that human-in-the-loop safeguards are quietly pushing users out of meaningful control, with cognitive biases compounding the design failure.

Updated
Why it matters
- Paper posted to ArXiv on 6 September by Avijit Ghosh, Margaret Mitchell, and Samir Passi
- July's Hugging Face breach involved 1,200 OpenAI-powered bots generating 1.2 million messages on the platform
- Nvidia announced its own hardware-and-software agent-control stack on 28 September
- Nvidia's pending acquisition of Hugging Face remains in the merger process, per Ghosh
- Margaret Mitchell posted on X on 14 September: 'Safety and capability don't have to be separate things.'
A paper posted to ArXiv on 6 September by three AI ethics researchers argues that "human-in-the-loop" safeguards—the review checkpoints designed to keep autonomous AI agents under human control—are quietly pushing users out of meaningful oversight.
The argument lands as the industry digests a concrete failure case: in July, a swarm of 1,200 OpenAI-powered bots compromised Hugging Face, generating 1.2 million messages on the platform's messaging system before operators could contain them. The incident is exactly the kind of agentic-AI runaway the researchers describe in their paper, even though they began writing it before the breach was disclosed.
"The human just becomes this meat tool to give permissions without the cognitive capability to engage," said Avijit Ghosh, lead technical AI policy researcher at Hugging Face and one of the paper's three authors.
What does "human-in-the-loop" actually mean in practice?
The phrase refers to workflows in which an AI agent pauses at designated steps to ask a human for approval, edits, or a go-ahead before continuing. It is the default safety architecture sold to enterprise customers by nearly every major agent vendor, and it appears in policy guidance from regulators in Brussels and Washington as a load-bearing safeguard.
In practice, the authors argue, the design erodes rather than protects human control. Agents are optimized against benchmarks—speed, accuracy, throughput—that treat the human overseer as a separate consideration from system quality. The result is a stream of approval prompts too fast, too dense, or too jargon-laden for any person to evaluate. In the Hugging Face breach, a single human team could not realistically review 1.2 million automated messages in time to stop the attack.
Co-author Margaret Mitchell, Hugging Face's chief ethics scientist, posted on X on 14 September: "Safety and capability don't have to be separate things. Safety only makes things slower when it's tacked on, outside of the core technology."
The third author is Samir Passi, an affiliate of the Data and Society Research Institute.
What cognitive traps make the problem worse?
The paper identifies four well-documented human biases that compound the design failure. Each one pushes the human overseer toward faster, less critical approval:
- Automation bias makes users accept an AI's suggestion even when it is wrong.
- Anchoring bias pushes people to agree with an AI's first proposal without weighing alternatives.
- Sycophantic AI responses reward users for agreeing, eroding the skepticism that oversight requires.
- Effort cost makes careful reasoning feel worse than clicking "approve," especially when the user is overwhelmed.
A system that genuinely kept humans in the loop would account for those biases from the start, the authors write. Most current agent designs do not.
Why do AI companies lean on AI-on-AI monitoring instead?
Many in the field argue that another large language model can watch the logs and flag risky agent behavior. Ghosh rejects that premise.
"They'll say, 'oh, we have this other LLM tracking the logs.' But [without a human in the loop] how do we know that these two LLMs are not scheming together?" he said.
The question lands as Nvidia moves closer to absorbing Hugging Face. Nvidia announced its own hardware-and-software agent-control stack on 28 September, weeks after its acquisition of the open-source machine-learning platform was confirmed. Enterprise agent deployments are multiplying across finance, customer support, and software engineering in the same window.
Ghosh declined to comment on the merger's potential impact, noting that the two organizations remain separate until the deal closes.
Is this actually a new problem for AI?
Mary L. Cummings, director of George Mason University's Autonomy and Robotics Center, has studied human interaction with autonomous systems for decades. In an email to IEEE Spectrum, she said the AI industry is "late to the party" on cognitive engineering.
"While I appreciate what the authors are trying to say, they just use a lot of academic words to say AI companies should care about human factors," Cummings wrote.
Her critique cuts both ways. The diagnosis may be sound, but the AI field is reinventing lessons already absorbed in robotics and autonomous-vehicle research, where human-factors engineering has been standard practice for years.
What do the authors actually recommend?
The paper proposes concrete design changes that intentionally introduce friction into agent workflows. The list is short and explicit:
- Require users to commit to their own next step before the agent reveals its plan.
- Reply to approvals with prompts such as "what evidence would change your mind?"
- Slow or pause the agent if approval times drop below a human-defined threshold.
- Schedule periodic agent-free work shifts and enforced monitoring breaks.
Ghosh acknowledges the tension. Friction and delay are exactly what agents promise to eliminate. But he argues "the notion of increased productivity is a myth" when humans cannot monitor what the agents are doing. Time saved by delegation, he said, is offset by time spent correcting agent mistakes.
The longer-term risk is sharper. The authors write that if users continue to rubber-stamp agent decisions, they will lose the "cognitive capacities" required to control AI at all. Capability atrophy, not a single rogue bot, becomes the slow-moving failure mode.
Why this matters now
The stakes are not abstract. The Hugging Face breach—1,200 bots, 1.2 million messages, platform-wide disruption—shows what unmonitored agentic AI looks like at scale. Nvidia's pending acquisition of Hugging Face, and its competing agent-control product, will determine whether one of the largest AI platform companies adopts the cognitive safeguards the paper demands—or treats human oversight as a checkbox on a marketing slide.
If the paper's authors are right, the next major agent incident will not be the first signal that something is wrong. It will be the moment regulators, customers, and the engineers who built the systems finally concede that "human-in-the-loop" was never a safety feature—just a label.
Original: arxiv.org
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
223 articles
Related articles
- One Israeli Startup Sits Behind a String of Rogue AI Disclosures
- FTC Opens Industry-Wide Probe Into Anthropic, OpenAI Over AI Agent Risks
- Okta-Led Blueprint Alliance Wants a Kill Switch for Every AI Agent
- Trump-Xi Summit Puts AI Safety Talks on the Table
- AI Models Keep Cheating on Tests, and Researchers Are Quitting