DNA helix as an abstract tapestry of life

(© Christian - stock.adobe.com)

In a Nutshell

  • Evo 2 predicts whether a genetic mutation is harmful with no task-specific training, beating other AI tools on several key tests.
  • It generates realistic genome-scale DNA from scratch, including sequences that mimic human, bacterial, and yeast genomes.
  • It is fully open source and shipped with safety guardrails to cut the risk of misuse.

A new artificial intelligence model taught itself to read and write DNA, and it can now flag harmful mutations and build genomes from scratch.

Called Evo 2, the model came out of a large team led by scientists at the Arc Institute, Stanford University, and NVIDIA. It was trained on nearly 9 trillion curated DNA base pairs drawn from bacteria, single-celled organisms called archaea, eukaryotes, and bacteriophages. That range lets one AI system work across virtually every known category of life on Earth. The primary research appeared in the journal Nature; this article summarizes a commentary by Guang-Guo Ying in AI & Environment, which reviewed that work.

What sets Evo 2 apart from earlier tools is what it manages without explicit instructions. Much like a person who reads well enough to spot a typo without being taught grammar rules one by one, Evo 2 has absorbed the underlying logic of DNA so thoroughly that it can flag dangerous genetic mutations, generate realistic genome-length sequences, and pick out features tied to biological concepts such as gene boundaries and regulatory signals, all on its own.

Teaching Evo 2 the Grammar of DNA

To see why Evo 2 matters, it helps to treat DNA as a language, one with billions of letters that spell out instructions for building and running every living thing. Most biology AI tools have learned only one dialect, focusing on either bacteria or humans, but not both. Evo 2 trained on the full library.

At the core of its design is a context window, basically how much of the DNA text the model can take in at once. Evo 2 can process up to 1 million genetic letters at a time, the longest reach of any DNA-focused AI model so far. That matters because many of the most telling patterns in DNA do not show up in short stretches. They only make sense across large sections read together, the way a plot twist only lands for someone who has followed the story from the start.

Evo 2 comes in two sizes, one with 7 billion parameters and one with 40 billion, and runs on a new architecture that makes the million-letter window possible.

How Evo 2 Spots Dangerous Mutations

One of the sharpest things Evo 2 can do is judge genetic mutations, the small DNA changes that can cause disease, with no prior coaching on what a harmful mutation looks like. Researchers call this zero-shot prediction, meaning the model relies entirely on what it picked up during training, with no extra guidance.

Tested against ClinVar, a large database of human genetic variants that doctors have studied, Evo 2 beat competing unsupervised methods for mutations that sit outside protein-coding regions, stretches that have long been harder to interpret. It also ranked first among zero-shot models on a separate test focused on mutations that change how genes get read and processed.

BRCA1, a gene widely known for its link to breast and ovarian cancer, offered a clear example. Evo 2 told harmful mutations apart from harmless ones with high accuracy. When its internal data trained a lightweight secondary classifier, that system reached an area under the precision-recall curve of 0.88, well ahead of simpler baseline approaches.

Infographic showing how Evo 2 learns from nearly 9 trillion DNA base pairs to predict genetic variants, generate genome-scale DNA sequences, and design regulatory DNA that was tested in laboratory mouse and human cells.
Infographic by StudyFinds
Writing New Genomes From Scratch With Evo 2

Beyond reading DNA, Evo 2 can write it. Given a short DNA sequence as a starting prompt, the model can produce entire genes or even genome-length sequences that look biologically realistic by several measures.

Prompted with a segment of the human mitochondrial genome, the DNA inside the energy-producing parts of cells, Evo 2 generated sequences roughly 16,000 letters long that held the expected protein-coding genes, tRNAs, and rRNAs in the right order. Given a piece of Mycoplasma genitalium, it produced nearly complete genomes close to 580,000 letters long, with about 70% of the predicted genes closely matching known protein families, a step up from earlier models. It also generated realistic yeast DNA, matching expected patterns for gene and tRNA density.

Researchers also put Evo 2 to work designing DNA meant to produce specific chromatin accessibility patterns, which control how open or closed a stretch of DNA is inside a cell’s nucleus and help decide whether nearby genes switch on or off. Pairing Evo 2 with a set of guiding prediction tools, the team designed sequences several thousand letters long, synthesized them, and inserted them into real mouse and human cells. In mouse embryonic stem cells, the designs hit the intended patterns with AUROC scores between 0.92 and 0.95. In human cell lines, more than 90% of the designs produced the wanted result, whether the goal was a pattern unique to one cell type or shared across two.

One caveat: these generated sequences have not yet been shown to work inside living cells, which the researchers call an important next step. Even so, their structural realism across such different organisms stands as an achievement on its own.

Built to Be Shared, and Built to Be Safe

Evo 2 and all of its training data are fully open source, so any researcher anywhere can use them. Safety came built in. Sequences from eukaryote-infecting viruses were kept out of the training data, and tests confirm the model does poorly when asked to predict or generate human viral sequences, which lowers the risk of accidental misuse.

Fairness got attention too. Researchers checked whether the model’s predictions might work against people of non-European ancestry, a known problem with many existing genetic tools, and found no systematic disadvantage.

Evo 2 is a direct successor to Evo 1, an earlier model trained only on bacterial genomes. Going from one to the other marks a big jump in both scope and reach, from a tool locked to a single branch of life to one that spans microbes, mammals, and everything between.

Authors of the underlying research call Evo 2 a major step toward a single model of biology, one that can read, interpret, and write genomic blueprints across the full range of living things. For a field long stocked with narrow tools built for narrow problems, one model fluent in DNA across the whole tree of life changes what researchers can attempt.


Paper Notes

Limitations

Ying and the underlying study are upfront about a few caveats. Evo 2’s generated DNA looks biologically realistic by several measures, but whether those sequences would actually function inside living cells has not been shown. Running the model’s most advanced guided-generation methods costs a lot of computing power, which could put them out of reach for some labs. Folding in real experimental feedback, through fine-tuning or reinforcement learning, might improve both speed and accuracy, though that work has not been done yet.

Funding and Disclosures

Guang-Guo Ying, who wrote the AI & Environment commentary this article draws on, reports that no funding was received to prepare the manuscript and declares no conflict of interest.

Publication Details

Author: Guang-Guo Ying, School of Environment & Environmental Research Institute, South China Normal University, Guangzhou, Guangdong, China.

Article Title: “Evo 2 writes the book of life across all domains.”

Journal: AI & Environment (AI Environ.), 2026, Volume 1, Issue 2, pages 55–56

DOI: 10.66178/aie-0026-0010

Dates: Received April 16, 2026; Revised April 16, 2026; Accepted May 22, 2026; Published Online June 8, 2026. Primary research referenced: Brixi G, Durrant MG, Ku J, Naghipourfar M, Poli M, et al. 2026. Genome modelling and design across all domains of life with Evo 2. Nature 652:1349–1361 . Published open access under a Creative Commons Attribution 4.0 (CC BY 4.0) license by New Horizon Press Limited.

About StudyFinds Analysis

Called "brilliant," "fantastic," and "spot on" by scientists and researchers, our acclaimed StudyFinds Analysis articles are created using an exclusive AI-based model with complete human oversight by the StudyFinds Editorial Team. For these articles, we use an unparalleled LLM process across multiple systems to analyze entire journal papers, extract data, and create accurate, accessible content. Our writing and editing team proofreads and polishes each and every article before publishing. With recent studies showing that artificial intelligence can interpret scientific research as well as (or even better) than field experts and specialists, StudyFinds was among the earliest to adopt and test this technology before approving its widespread use on our site. We stand by our practice and continuously update our processes to ensure the very highest level of accuracy. Read our AI Policy (link below) for more information.

Our Editorial Process

StudyFinds publishes digestible, agenda-free, transparent research summaries that are intended to inform the reader as well as stir civil, educated debate. We do not agree nor disagree with any of the studies we post, rather, we encourage our readers to debate the veracity of the findings themselves. All articles published on StudyFinds are vetted by our editors prior to publication and include links back to the source or corresponding journal article, if possible.

Our Editorial Team

Steve Fink

Editor-in-Chief

John Anderer

Associate Editor

Leave a Comment