Imagine you are deciding whether to put money into a startup, and someone hands you a research report. At the top, a line explains how it was made. One version says a financial analyst wrote it. Another says an AI system produced it. A third says an AI produced it “under the supervision of” a licensed analyst. Which report would you trust most?
The popular assumption, shared by regulators, tech companies, and much of the academic literature, is that the version with a human watching over the AI should win. Adding a person to the pipeline is widely treated as the way to make AI trustworthy. A new study offers evidence that this assumption is incomplete, and sometimes backward.
The question behind the trust
The research, published in Psychology & Marketing, was conducted by Guneet Kaur Nagpal and June Cotte of the Ivey Business School at Western University in Ontario. Their starting point was a gap in what researchers actually know. Studies repeatedly show that human involvement tends to make AI-generated content feel more credible, but they rarely explain why. What exactly does a “human in the loop” signal to a reader?
The authors argue that a credibility judgment is not really a reaction to the source itself. Instead, a label like “written by an analyst” or “produced by AI” acts as a mental shortcut. It invites the reader to fill in a story about how the work was made and whether it can be trusted. The team wanted to identify which specific story dominates.
They laid out five candidate explanations that a phrase like “human-supervised” might trigger. A reader might infer that a human expert is simply smarter than a machine (competence). Or that someone can be reached to explain the work (answerability). Or that someone can be blamed if it goes wrong (blame traceability). Or that a qualified person actually checked the work before release (what the authors call oversight adequacy). A fifth possibility runs the other way: mixing two parties might spread responsibility so thin that no one seems clearly in charge, which could reduce trust rather than build it.
The researchers theorized that one of these, oversight adequacy, does most of the work. Their reasoning was that only this signal directly answers the question a credibility judgment is really asking: was this report done well? Knowing that someone could be contacted, or is generally capable, does not tell you whether anyone actually reviewed this particular piece of work.
Testing four kinds of AI report
In the first experiment, the team recruited 620 U.S. residents through the research platform Prolific, all with at least two years of personal investing experience. After exclusions, 570 remained. Everyone read the exact same one-page investment report about a real crowdfunding company. The only thing that changed was a short intro paragraph describing how the report was prepared.
Participants were randomly sorted into four versions. In one, the report came from a named human analyst. In another, from a “black-box” AI, meaning a system whose inner workings are hidden. A third credited an AI supervised by a human. The fourth described a “glass-box” AI, one that disclosed its reasoning, including which factors it weighed, what model it used, and where its data came from. Readers then rated the report on a 16-item credibility scale and made a mock investment decision.
The labels moved credibility only slightly, and did not change investment behavior at all. Two patterns stood out. The glass-box AI, with no human involved, scored about the same as the human-written report. And the human-supervised AI, the exact setup that European and U.S. regulators treat as the trustworthy default, scored the lowest of all four versions on every measure.
The researchers interpret this as a case where the vague phrase “human-supervised” backfires. When the human’s role is left undefined, readers may sense that responsibility is spread across parties without gaining any assurance that real checking happened.
Pinpointing the signal that matters
The first study raised a puzzle but could not explain it, because it never measured the five psychological inferences directly. The second experiment did. This time the team recruited investors from a North American business school lab, ending with 296 participants in a design that crossed two factors. One factor varied whether the report came from a human alone or from an AI working under a human. The other switched a formal accountability disclosure on or off. When on, the report named a responsible analyst, gave a contact email, and included a governance statement. When off, it named only the issuing firm.
Participants filled out a new 12-item scale measuring four of the candidate signals, so the researchers could see which inference carried the credibility effect. In statistics, when one variable helps explain the link between two others, it is called a mediator. The team wanted to know which signal mediated the path from source label to credibility.
Perceiving human involvement raised credibility and eight related outcomes by roughly half a standard deviation, a moderate effect. And the analysis showed that perceived oversight adequacy accounted for about 84 percent of that link. Once the researchers factored in whether readers believed the work had been competently reviewed, the direct effect of the source label essentially vanished. The other signals carried far less weight. This pattern held all the way through to behavior: the sense that the work was reviewed fed into credibility, which in turn fed into readers’ intent to follow the recommendation, share the report, trust it, and rely on it.
There was also a piece of digging into the order of the reasoning. The team found that readers seem to first decide whether the work was reviewed, and that judgment then shapes whether they see the producer as competent, rather than the reverse.
Naming who is responsible fell flat
The disclosure manipulation produced one of the study’s more pointed results. Naming a responsible analyst, providing a contact email, and adding a governance statement had no significant effect on credibility, on any of the accountability signals, or on investment decisions. This null result was well-powered, meaning the sample was large enough that a real effect of modest size would likely have shown up.
The authors distinguish between two things that are often bundled together. Formal accountability tells a reader who can be reached if something goes wrong. Perceived oversight adequacy tells a reader whether the work was actually checked. Their evidence suggests only the second one moves credibility. As they put it, disclosure “tells consumers who can be reached, but not whether the work was actually checked.”
What it may mean for firms and regulators
For companies putting out AI-assisted content, the study points toward a shift in tactics. Rather than stamping a report with a name and contact details, the authors suggest showing the process: what data was verified, how conclusions were stress-tested, what a reviewer actually examined. The glass-box result suggests that visible algorithmic reasoning can build the same credibility as a human reviewer when the process itself is inspectable.
The findings also sit uneasily with how current AI governance treats human oversight. The authors argue that disclosure requirements may still serve legitimate legal and enforcement purposes, but are unlikely on their own to build the consumer trust regulators seem to expect from them.
Several limits are worth keeping in mind. Mediation analysis rests on assumptions that cannot be fully tested, so the causal story remains an inference rather than a proven chain. The studies used a single investment report and two specific samples, one online and one a group of relatively young, less experienced student investors, neither representative of retail investors as a whole. The authors frame the generalization of their oversight-adequacy account to areas like medical or legal content as an open question for future work.
The takeaway they land on is a reframing. The useful question, they write, “is not whether to add a human, but how to signal that the analytical process was sound.”




