Therapy session

(Photo by Ilona Kozhevnikova on Shutterstock)

Why Two Psychiatrists Can Look At The Same Patient And See Two Different Illnesses

In A Nutshell

  • Two doctors reviewing the same written psychiatric case agreed on a diagnosis only about 55% of the time.
  • Schizophrenia cases were the least reliable, sometimes splitting almost evenly with OCD, a personality disorder, or bipolar disorder.
  • Experience made no real difference. Psychiatrists, residents, and other doctors agreed at similar rates regardless of years in practice.
  • Bipolar disorder, OCD, and depression cases were matched correctly 88% of the time or higher, showing the problem clusters around specific disorders rather than diagnosis in general.

A patient walks into one psychiatrist’s office and leaves with a diagnosis of schizophrenia. A different patient, describing textbook symptoms of the same disorder to a different doctor, walks out labeled with OCD, a personality disorder, or bipolar disorder instead. That’s what a massive new international study found when it tested more than a thousand doctors using identical written patient cases.

Published in Frontiers in Psychiatry, the study gathered 1,038 medical doctors working in adult psychiatry across 19 countries in Europe and South America. Each one read written descriptions of fictional patients and assigned a diagnosis. Because everyone read the same words describing the same patients, disagreement couldn’t be blamed on one doctor asking different questions than another, or a patient having a rough day. Judgment alone was the variable, and it turns out to be wildly inconsistent when it comes to certain mental illnesses.

Two doctors looking at the same case agreed on a diagnosis only about 55% of the time. Doctors landed on the reference diagnosis established by an expert panel just 66% of the time overall. For cases involving schizophrenia, that figure sank much lower, with doctors frequently mistaking it for OCD, a personality disorder, or bipolar disorder, depending on which version of the case they saw.

Nine Written Cases, 1,038 Doctors, Two Diagnoses Each

Researchers built nine fictional patient cases reflecting disorders psychiatrists see regularly: three variations of schizophrenia, plus single cases representing schizotypal disorder, bipolar disorder, depression, generalized anxiety disorder, OCD, and a personality disorder. Each doctor was randomly given two of the nine cases to diagnose, choosing from a list of 30 common mental health diagnoses, producing 1,902 diagnostic decisions total. Most participants were psychiatrists, though psychiatric residents and other doctors working in adult psychiatry took part too, with experience ranging from brand-new doctors to those with 15 or more years in the field.

Experience made almost no difference to how often doctors agreed with each other. A doctor practicing for two years fared about the same as one with twenty, and psychiatrists matched residents and other medical doctors at similar rates. Job title and years on the job showed no meaningful link to agreement, pointing away from a simple lack of training. Something else drives the split.

That something else appears to be which disorder is on the table. Some cases produced near-unanimous agreement: bipolar disorder matched the reference diagnosis for 94% of doctors, OCD for 92%, depression for 88%. But schizophrenia cases told a different story. Across the study’s three schizophrenia vignettes, doctors often split between schizophrenia and alternatives, with one case dividing almost evenly between a personality disorder (45%) and schizophrenia (43%), and another seeing 41% call it OCD while only 37% matched schizophrenia.

psychiatry diagnoses infographic
New research shows psychiatric diagnoses vary widely between doctors, even under standardized conditions. (Image by StudyFinds)

Wrong Diagnoses Clustered Around the Same Alternatives

Wrong answers weren’t scattered randomly. When doctors missed a schizophrenia diagnosis, they tended to land on the same alternatives again and again, mainly OCD, a personality disorder, or bipolar disorder. Study authors argue doctors weren’t confused or guessing blindly. Instead, the pattern suggests they were weighing the same symptoms differently, often anchoring on obvious features like obsessive behaviors while giving less weight to the psychotic symptoms that should have pointed toward schizophrenia under standard classification rules.

Comparing these results to older research adds weight to the concern. The famous US-UK Diagnostic Project from the 1970s first exposed how differently American and British psychiatrists diagnosed the same patients, prompting a major overhaul of diagnostic manuals in the following decades. Those changes were supposed to fix the problem. This new study, along with agreement scores measured during the 2013 update to the American diagnostic manual, suggests it never fully took. Doctors here performed worse than doctors did during field trials for the third edition of the American diagnostic manual back in the 1980s, landing closer to the shakier scores seen in the decades before that manual existed.

A Wrong Label Can Follow a Patient Through Treatment

Diagnostic labels aren’t just paperwork. They shape treatment, what a doctor tells a family about the road ahead, and how a patient is understood by every future clinician who reads their chart. If similar disagreement plays out in real practice, it could affect treatment decisions, prognosis, and continuity of care.

It matters for research too. Scientific studies of mental illness assume a group of “schizophrenia patients” all share the same underlying condition. If doctors sort patients into that category inconsistently, research samples may be muddier than assumed, which could help explain why it’s been hard to find clean biological explanations or reliably effective treatments tied to specific psychiatric labels. Study authors note this may also explain why transdiagnostic approaches, treatments that don’t rely on a single precise label, have been gaining ground.

Real limits apply here, though. Written cases are a controlled, near best-case scenario compared to a live interview, where confusion, incomplete information, and rushed appointments can make things messier still. Researchers themselves suggest their agreement numbers likely mark a ceiling, meaning ordinary practice could produce even lower agreement than what showed up here.

Fifty years after the field’s reliability problem was first exposed, and decades after supposedly fixing it with stricter diagnostic rulebooks, psychiatry still hasn’t closed the gap between two doctors looking at the same case and reaching the same conclusion. That’s a crack running through how mental illness gets identified, treated, and studied.


Disclaimer: This article summarizes findings from a peer-reviewed study for general informational purposes and is not a substitute for professional medical or psychiatric evaluation. Anyone with concerns about a diagnosis should speak with a licensed mental health provider.


Paper Notes

Limitations

Written case vignettes rather than live patient interviews are a core limitation the authors acknowledge, since vignettes cannot capture the full complexity of real clinical encounters, including dynamic conversation, patient disclosure, and the ability to ask clarifying questions. Participation was voluntary and recruited through professional networks, so response rates could not be calculated, and clinicians with a stronger interest in diagnosis or academic psychiatry may have been more likely to take part, potentially making the reported agreement look better than it would in a broader population of doctors. Researchers grouped some related diagnoses into broader categories and limited diagnostic choices to a list of 30 commonly used disorders rather than the full range available under the ICD-10 classification system, both of which likely increased the measured agreement compared to what might be seen with full diagnostic granularity. Authors also note that doctors were not explicitly instructed to apply formal diagnostic criteria step by step, so results may reflect how doctors diagnose in routine practice rather than a strictly rule-based procedure. Additionally, providing a predefined list of diagnostic options may have shaped participants’ decisions by prompting them to consider options they might not have otherwise thought of.

Funding and Disclosures

One author, Dr. Boberg, was supported by a grant from Region Zealand, Denmark. The funder had no role in the study’s design, data collection, analysis, interpretation, writing, or the decision to submit the paper for publication. Several authors reported financial relationships with pharmaceutical or industry entities within the past 36 months, including one author who received payment for lectures, presentations, and travel support from Johnson & Johnson, Lundbeck, and Bial, and another who received travel support from Acadia. Two authors disclosed that they were editorial board members of Frontiers at the time of submission, which the journal states had no impact on the peer review process or final decision. The authors stated that generative AI was not used in the creation of the manuscript.

Publication Details

Titled “Reliability of psychiatric diagnoses in the 21st century,” the paper was authored by Mateo Boberg, Mads Gram Henriksen, Jonas Berge, Nikolas Fascendini, Martin Jandl, Hanna Karakuła-Juchnowicz, Luís Madeira, Daniele Rossi Grauenfels, Hedda Soloey-Nilsen, Ana Giurgiuca, Igor Studart, Guilherme Messas, Tamara Ojeda Uribe, Luis Varela, Ricardo Corral, Michel Cermolacce, Tudi Gozé, Léon Franzen, Stefan Borgwardt, Jeff Huarcaya-Victoria, Konstantina Karkala, Ramune Mazaliauskiene, Jonas Eberhard, Troels Schmidt, Jean-Arthur Micoulaud-Franchi, Borut Skodlar, and Julie Nordgaard. It was published in Frontiers in Psychiatry on July 29, 2026. DOI: 10.3389/fpsyt.2026.1891625.

About StudyFinds Analysis

Called "brilliant," "fantastic," and "spot on" by scientists and researchers, our acclaimed StudyFinds Analysis articles are created using an exclusive AI-based model with complete human oversight by the StudyFinds Editorial Team. For these articles, we use an unparalleled LLM process across multiple systems to analyze entire journal papers, extract data, and create accurate, accessible content. Our writing and editing team proofreads and polishes each and every article before publishing. With recent studies showing that artificial intelligence can interpret scientific research as well as (or even better) than field experts and specialists, StudyFinds was among the earliest to adopt and test this technology before approving its widespread use on our site. We stand by our practice and continuously update our processes to ensure the very highest level of accuracy. Read our AI Policy (link below) for more information.

Our Editorial Process

StudyFinds publishes digestible, agenda-free, transparent research summaries that are intended to inform the reader as well as stir civil, educated debate. We do not agree nor disagree with any of the studies we post, rather, we encourage our readers to debate the veracity of the findings themselves. All articles published on StudyFinds are vetted by our editors prior to publication and include links back to the source or corresponding journal article, if possible.

Our Editorial Team

Steve Fink

Editor-in-Chief

John Anderer

Associate Editor