(Credit: © Sandwish - stock.adobe.com)
Every AI Chatbot Tested Offered Less Variety Than a Google Search, Study Finds
In A Nutshell
- Larger AI models tended to give narrower, more repetitive answers than smaller versions from the same model family.
- Every one of the 27 systems tested, even the newest, produced less varied information than top Google results.
- Variety in AI answers has improved in recent years, and letting AI pull from current web pages helped, though unevenly.
- For five of eight countries tested, AI answers reflected English-language knowledge more than local-language knowledge.
Ask a chatbot the same question enough different ways, and the biggest AI models start doing something strange: they hand back less variety, not more. A new study covering 27 popular AI systems and 70 million individual statements found that larger language models tended to produce narrower, more repetitive information than smaller versions of the same models, even though bigger models are generally built to perform better.
That finding lands at an awkward moment. Search engines are increasingly steering people toward AI-written summaries instead of a list of links, and there is speculation that people will soon get most of their information through an AI middleman. Researchers from universities in Denmark and the United States wanted to know what happens to the diversity of information when that middleman does the talking. Their paper, posted to arXiv, an open site where researchers share papers, tests a concern known as “knowledge collapse,” where the range of ideas people can access shrinks into a narrow set of views.
Results came back mixed. Diversity in AI answers has improved over the past few years, which pushes back against the gloomier predictions. But every system tested, no matter how new or advanced, still produced less varied information than simply reading the top results of a Google search.
Researchers Measured AI Diversity Across 27 Models and 70 Million Claims
To measure how much variety shows up in AI answers, the research team built what amounts to a giant claim-sorting assembly line. They tested AI systems from four families, Llama, Gemma, Qwen, and OpenAI’s GPT models, ranging from small systems with about a billion internal settings to large ones with more than 70 billion. Each model answered questions on 155 topics, from broad subjects like democracy, free speech, and nuclear weapons to history and public figures tied to 12 countries, including the United States, India, Russia, Brazil, and Saudi Arabia. Country-specific picks ranged from the Berlin Wall to K-pop.
Every model answered each topic using 200 different human-written prompt styles, such as “write me an essay about” a given subject. That produced 1.7 million individual responses. An AI model then broke each response into individual factual statements, and human reviewers checked samples of that work. Statements making the same basic point were grouped together, so sentences all saying, for example, that democracy involves free elections counted as one core piece of information instead of many.
Diversity was scored with a formula ecologists use to count species in a habitat and how evenly the population is spread among them. Applied here, it rewards a model for covering many distinct points rather than repeating the same few. As a point of comparison, the researchers also pulled the top 40 Google search results for each topic.
Every AI Model Tested Trailed Web Search on Diversity
Not every model family improved, with Qwen the main exception. Google’s Gemma 3 and OpenAI’s GPT-5 posted the biggest jumps between generations, a sign that recent updates from those companies are paying off. On the study’s topics, GPT-5 offered more than five times the variety of an earlier OpenAI model, GPT-4o.
So how big is the gap? Search results were at least 18.7 percent more varied than the best-performing model, GPT-5. Because the researchers built the search comparison to be conservative, the real gap is probably wider.
One technique helped. Called retrieval-augmented generation, it lets an AI pull information from current web pages before answering instead of relying only on what it absorbed during training. Answers got noticeably more varied, but the boost was uneven. Topics tied to the United States, India, Russia, France, and China gained the most, while other countries gained less. The researchers think their U.S.-based search tool may have under-represented information specific to other countries.
Larger AI Models Produced Less Varied Answers Than Smaller Ones
Model size produced one of the odder results. After accounting for other factors, larger models were linked to less diversity, meaning smaller versions within the same model family tended to generate more varied claims than their bigger siblings. The researchers have a possible explanation: bigger models may memorize more of their training data, which could narrow what they say when asked about the same topic again and again.
AI Answers Leaned on English-Language Knowledge in Five of Eight Countries
A separate test asked whose knowledge shows up in AI answers. The researchers matched AI claims about country-specific history and public figures against English-language Wikipedia pages and against the same topics in each country’s own language, using eight countries. Prompts were written in English. For five of the eight countries, AI answers reflected the English-language pages clearly more than the local-language ones, and in no country did local-language material come out clearly ahead. Topics tied to the United States were represented more strongly than topics tied to any other nation.
Progress is real, but it has not reached everyone equally. As search increasingly hands people AI-written overviews, it matters that AI closes the gap with web search and gives local knowledge a fairer share of the answers. Otherwise, a tool meant to broaden what people know could narrow it without anyone noticing.
Paper Notes
Limitations
Researchers note that their topic selection combined random sampling with hand-curated choices meant to cover general concepts, historical events, and public figures across 12 countries, which may limit how broadly the results generalize. Their approach to measuring “whose” knowledge is represented worked at the level of entire countries rather than more specific communities, and for some countries the topic selection was made from an outside perspective. The English-versus-local-language comparison covered eight countries and used English prompts, so it does not speak to every non-English-speaking country. Topics were limited to those whose English Wikipedia pages met a minimum content quality rating, which excluded very obscure or niche subjects that might behave differently. The team used 200 prompt templates due to computational limits and acknowledges that broader sampling could add more coverage. Their process for breaking responses into claims and grouping them, while validated by human reviewers, is not perfect and introduces some unavoidable noise into results. Their retrieval-augmented setup was designed to simulate a realistic scenario, but real-world retrieval-augmented systems will vary in how they are built. The memorization explanation is the authors’ hypothesis rather than a tested result, and the authors add that diversity should be weighed alongside other qualities of knowledge, such as factuality and relevance. This work measures how varied AI-generated claims are compared with web search, and does not show that knowledge collapse is already happening across society.
Funding and Disclosures
Acknowledgments in the paper note that one author was supported by a Danish Data Science Academy postdoctoral fellowship, another by the Stanford Interdisciplinary Graduate Fellowship, the Stanford Center for Affective Science Graduate Fellowship, and the Future of Life Institute Vitalik Buterin PhD Fellowship, and three other authors were supported by the Pioneer Centre for AI under a Danish National Research Foundation grant. The authors also disclose making minimal use of the AI assistant Claude in the final polishing stages of preparing the manuscript, primarily for feedback on figures, consistency of claims, and framing.
Publication Details
This paper is titled “What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models,” authored by Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park, and Isabelle Augenstein, affiliated with Aalborg University Copenhagen, the University of Copenhagen, Stanford University, the University of Colorado Boulder, and the University of Texas at Austin. It was posted to the preprint server arXiv (computer science, computation and language category) as arXiv:2510.04226 (version 7, dated August 31, 2026), DOI 10.48550/arXiv.2510.04226, available at https://arxiv.org/abs/2510.04226. Code and data are hosted at https://github.com/dwright37/llm-knowledge.







