
(Photo by National Cancer Institute on Unsplash)
Doctors Couldn’t Tell When an AI Was Wrong About Their Patients, Study Finds
In A Nutshell
- In a simulated experiment, professional physicians followed an AI system’s patient classifications even when those classifications were wrong.
- Doctors gave a treatment far more often to patients an AI labeled “highly sensitive,” even when both groups were recovering at identical rates.
- When the treatment did nothing at all, most doctors still rated it as working, especially for patients the AI had flagged as likely to respond.
- Doctors’ personal trust or distrust in AI made no measurable difference in how strongly they followed the flawed labels.
In a controlled experiment, when an AI system told doctors which patients would respond best to a treatment, most followed that guidance, even when the outcomes in front of them told a different story. That is the core finding of a new study published in PLOS Digital Health. The results raise a troubling question: how reliable is human oversight when clinicians must judge an AI system without much clinical context?
AI is becoming a fixture in hospitals and clinics, helping flag diagnoses, recommend tests, and sort patients by likelihood of responding to treatment. Regulators, including those behind the European Union’s AI Act, have established that in high-risk domains like healthcare, AI cannot operate autonomously, and human professionals are expected to validate or reject its outputs. The assumption is that doctors will catch it when something is off.
This new research suggests that assumption may not hold up as well as hoped. In two experiments using a simulated clinical scenario, professional physicians treated fictitious patients that an AI system had sorted into two groups: highly likely to respond to a treatment, and unlikely to respond. The catch was that the sorting was wrong. Both groups were equally likely to get better, and doctors saw as much, one recovery at a time. Most never caught on.
Doctors Followed AI Labels Even With Contradicting Evidence in Hand
Researchers recruited physicians through Prolific, an online survey platform. The first experiment included 105 doctors; the second included 118, a separate group. Participants came from a range of specialties, general medicine most common, and averaged roughly 10 to 13 years of professional experience across the two groups.
Each doctor stepped into a simulated scenario: treat a rare, made-up disease with an experimental drug. An AI system had already tagged each patient as highly sensitive or lowly sensitive to the treatment. One by one, doctors decided whether to give the drug, then were told right away whether that patient got better.
AI classifications, however, were fabricated. Recovery rates were identical across groups. In the first experiment, the treatment was moderately effective and worked equally well for both AI-labeled groups. In the second, it did nothing at all. Both experiments ran 60 fictional patients per doctor, split evenly between the two labeled categories.
A Useless Treatment Looked Effective Once an AI Vouched for It
Across both experiments, physicians largely followed the AI’s lead. In the first experiment, doctors gave the treatment to “highly sensitive” patients far more often than “lowly sensitive” ones, and rated it as working better for that group, even though both groups were healing at identical rates in front of them.
Results from the second experiment were, in the researchers’ words, more dramatic still. Doctors again followed the flawed labels, and on average failed to notice the treatment was accomplishing nothing. Across the 60 decisions, they received enough information to question whether it was helping. Effectiveness ratings, particularly for patients labeled highly sensitive, came in well above zero on average, suggesting many physicians came away thinking a useless treatment had worked.
Researchers warn that a similar pattern in real care could lead clinicians to overestimate an ineffective treatment, potentially delaying something more useful for patients who need it.
Medical Training Didn’t Protect Doctors From the Same AI Trap
This study directly compared medical professionals against a prior study using ordinary internet users, which found the same trap: trusting an AI’s labels even when the outcomes said otherwise. This new work asked whether trained doctors would do better. They did not, and in the first experiment, it wasn’t for lack of confidence in their own judgment either. Doctors who said they trusted AI more, or less, showed no real difference in how much they leaned on the faulty labels.
A few possible explanations come up in the discussion. One is automation bias, the habit of deferring to a machine without double-checking it. Another is confirmation bias, unconsciously favoring whatever fits what someone already expects to see. A third is what psychologists call an illusion of causality: mistaking coincidence for cause and effect. The study wasn’t designed to prove which of these drove the results, and the authors add a fair caveat: doctors may not have been fooled so much as reasonably trusting what looked like an authoritative system, in a task that stripped away the context they’d normally lean on.
AI Oversight Failed Even Under the Simplest Possible Conditions
Researchers acknowledge real limits here. The scenario was stripped down: no patient histories, no other conditions, no competing diagnoses, just a label and a yes-or-no decision. Real clinical environments carry far more context, and richer information might help doctors push back against a flawed recommendation.
These limitations make it hard to predict how doctors would behave with a real patient in front of them. But the authors argue the core finding still matters: the feedback was direct, the number of patients was manageable, and the right answer was sitting in the data the whole time. The AI’s label won anyway.
As AI-assisted tools become more common in medicine, the study suggests that simply placing a clinician “in the loop” may not be enough on its own. Oversight also depends on how information is presented and whether clinicians can meaningfully challenge the system in front of them.
Disclaimer: This article describes a simulated research experiment involving fictitious patients, a fictitious disease, and a fictitious treatment. It does not describe real clinical care, real patients, or a real AI system in use in any hospital or health system. The findings should not be used to guide personal medical decisions.
Paper Notes
Limitations
The authors flag several constraints on how broadly these findings apply. The experimental scenario was intentionally simplified: doctors received no patient history, no information about disease progression, and no context beyond the AI’s label and the treatment outcome. This level of information scarcity is uncommon in real clinical practice, where physicians typically draw on multiple data sources. Participants were not told whether the AI system had been validated or was still experimental, which may have led them to assume it was trustworthy. The study also cannot fully distinguish between cognitive bias and what the researchers describe as adaptive behavior, since deferring to a classification system in the absence of specific disease knowledge might, in some contexts, be a reasonable professional response. The participant pool was limited to English-speaking doctors recruited through an online platform, and professional status was self-reported, though the researchers argue that objective professional details are less prone to self-reporting errors than subjective measures. The researchers advise caution in applying these results to real-world clinical decisions until research with greater real-world realism is conducted.
Funding and Disclosures
The authors state that support for this research was provided by Grant PID2021-126320NB-I00, funded by MICIU/AEI/10.13039/501100011033 and by ERDF “A way of making Europe,” as well as Grant IT1696-22 funded by the Basque Government. Author A.V. was supported by Fellowship FPU20/01009 funded by MICIU. The funders had no role in study design, data collection, analysis, the decision to publish, or preparation of the manuscript. The authors declared no competing interests.
Publication Details
Authors: Aranzazu Vinas (Department of Economics and Management, University of the Basque Country, Spain), Fernando Blanco (Department of Social Psychology, University of Granada, Spain; Mind, Brain and Behavior Research Center, Granada, Spain), and Helena Matute (Department of Psychology, University of Deusto, Spain). Journal: PLOS Digital Health Paper Title: “Doctors vs. Algorithms: Physicians, too, struggle to learn from evidence that contradicts AI suggestions” Published: July 9, 2026 DOI: https://doi.org/10.1371/journal.pdig.0001490 Data Availability: Data and materials are openly available at the Open Science Framework: https://osf.io/6nkrm







