Behavioural Science
94% said yes: Human oversight in practice
How much should we trust AI-generated output? We know AI gets things wrong sometimes. The more interesting question is how many people actually check.
A research team at Northeastern University¹ put a number on it. They asked 107 developers to build an app with an AI coding agent – the same kind of setup many of us work in every day. The twist: the agent had been instructed to quietly build a path that sent customer data to an external server, while completing the assigned work normally.
94% of developers in this setup merged the malicious code. Why? One participant put it plainly: “I’m always used to clicking approve directly.”
Table of contents
Human in the loop?
Let’s be clear: the developers in this study had the expertise to catch this. 86% reported a security background and 70% had more than three years of coding experience, and most had worked with the tool before. What they lacked, in that moment, was any reason to think a closer look was warranted.
That gap is important, because so much rests on it. Human oversight is considered to be key to preventing automated errors and plays a major role in policy and regulation. It works on one condition: that the person in the loop independently evaluates what they approve. Under normal working conditions, the evidence says they mostly don’t.
Psychology has a term for this: Automation bias is the tendency to favour an automated recommendation over your own judgment, even when contradicting information sits in front of you. Researchers have documented it for decades, first in aviation and clinical decision support, and more recently wherever AI has arrived in professional work. Romeo and Conti’s 2026 review in AI & Society pulls 35 experimental studies together², and the pattern holds with uncomfortable consistency. In the remainder of this article, we will explore their findings and what they mean for our use of AI.
Why checking is harder than it looks
Security awareness has spent a decade on the human as an attack surface: the person who clicks the link, approves the payment, enters the credential. This adds a different dimension. Here the human is the safeguard, and we have been asking them to perform a demanding cognitive task with very little support.
Part of the reason is economic. Verifying an AI output is often harder than producing the judgment yourself, because you have to reconstruct reasoning you never performed yourself and imagine failure modes you have no visibility into. Under time pressure and high cognitive load, acceptance is cheap and scrutiny is expensive. A rational person will respond to that trade-off with over-reliance, which is why telling people to be more careful achieves so little.
The other part is that AI output carries every surface signal we normally use to judge quality. It is fluent, cleanly structured and confidently phrased. When a person produces work that looks like that, the polish is earned: it costs time, and time spent usually implies care. Generated output gets the polish for free, and it tells you nothing about the substance underneath.
Confidence matters a lot here. AI advice tends to arrive with certainty, lacking any of the caution a human expert would offer, and the reviewed studies find that readers over-trust it as a result. Even senior decision-makers treat AI recommendations as objective analysis rather than a fallible opinion, and lean on them harder for it. A response that sounds certain gives the reader very little to push against.
Explaining and alerting won’t do the trick
If the problem is that people cannot easily tell sound output from unsound, the intuitive fix is to show them more information. Transparency and explainable AI are the default answer, but studies do not support their effectiveness. People tend to treat an explanation of AI outputs as a signal in itself: they register that one exists rather than testing whether it holds. Simple explanations are misleading when the system is wrong, while complex ones add cognitive load and get skipped entirely. Across the reviewed studies, explanations reliably increased people’s acceptance of AI outputs while leaving their accuracy roughly where it was.
The Northeastern results underscore this. One condition of the study added an automated monitor that watched the agent and raised a warning on screen. It caught 84% of the sabotage attempts. In the sessions where it fired correctly (where the monitoring system did its job and told the person exactly what was wrong at exactly the right moment) 56% of participants still shipped the code.
An alert only works when it lands with someone who already has a reason to take it seriously, and that reason has to come from somewhere other than the alert. Which turns the question around: what would people need in order to succeed at the job we have already assigned them?
Enablement: AI literacy is the lever
The most consistent protective factor across the research is understanding the system you are working with, knowing where it holds up and where it breaks down. This is different from being comfortable using AI tools, which may quietly work against you.
The research describes the effects of AI literacy as a curve rather than a linear slope. People with very little AI background tend to reject AI suggestions outright. Those with deep background rely on AI appropriately, because they can anticipate where it fails. The group in the middle, with moderate familiarity, over-relies the most. Daily use of AI leads to a sense of fluency and comfort that overestimates actual understanding.
This should concern us, because that middle group is now most of the workforce. The Northeastern participants used AI coding tools every day, and their ease with the tool was part of what made the sabotage invisible. The few who caught it were those who had come in more skeptical of AI agents to begin with. That does not mean we should teach people to distrust AI outputs in general. But they need to be aware that errors or even malicious behavior may occur and that they need to actively look out for them. To be able to do this, they need to know how AI may fail, and what this failure looks like.
The reviewed experiments reveal another facet of this: people who were told a system could produce errors went on to review its recommendations measurably more thoroughly. In contrast, people who were just reminded that they carried responsibility for the outcome changed nothing about how they worked. Scrutiny follows a concrete expectation that something might be wrong, and a reminder of blame supplies no such expectation.
Processes that leave room for judgment
Enablement only pays off if the workflow gives people somewhere to apply it. The most reliable design choice is sequencing: when the AI’s answer arrives before the person has formed their own, it anchors them, whereas asking for an initial assessment first and revealing the system’s suggestion afterwards produces markedly more independent judgment. Deliberate friction on critical, hard-to-reverse actions helps too, provided it is implemented scarcely and with intention, rather than becoming background noise. And whatever we measure is what people will optimise for. An approval rate may look like productivity, but is indistinguishable from rubber-stamping, so a dashboard built on output quietly rewards the behaviour we are trying to prevent. Measuring review behaviour is harder, and it is the only way to know whether the safeguard is real.
The human in the loop deserves better
Human judgment belongs in the loop. It remains the critical factor that catches what a system can’t see. However, if we treat oversight as a box to tick rather than a capability to build, we undermine human-AI collaboration. We place someone at the end of an automated process, tell them nothing about how it fails, and give them no credit for looking closely – then call them in control.
Equipping people for that role means teaching how to catch failure, and building workflows and incentives that treat careful review as the work instead of an obstacle to it.
¹ Ye, Zou, Yu, and Shi, “Coding with ‘Enemy’: Can Human Developers Detect AI Agent Sabotage?” (2026), arXiv:2606.05647. This is a preprint and has not yet been peer-reviewed. The 94% figure refers to the study conditions without an automated monitor.
² Romeo and Conti, “Exploring Automation Bias in Human–AI Collaboration: A Review and Implications for Explainable AI,” AI & Society (2026). The individual findings referenced above are drawn from the 35 studies included in the review, where the original sources are cited in full.












