化学、生物与自动化实验室 颠覆级 有讲解视频
发表时间
2026-02-19
DOI
10.1038/s41586-026-10176-5

收录解读

问题与背景:这篇论文试图把基因组建模从局部任务模型提升为跨生命全域的统一基础模型。传统基因组模型通常只覆盖特定物种、特定长度或特定任务,而 Evo 2 的目标是同时覆盖细菌、古菌和真核生物序列,并把预测与设计放进同一框架。

方法/新意:论文提出 Evo 2 这一大规模基因组 foundation model,使用极长上下文的序列建模策略,在统一语料上学习跨物种、跨功能层级的表示。它不仅用于序列补全、功能推断和变异效应预测,也支持生成式设计,使模型从“读基因组”扩展到“写基因组”。

意义/放在仓库中的位置:这是 AI-enabled genomics 主线里的高位条目,和 AlphaGenome 同处“基因组基础模型”方向,但更强调跨生命全域与生成设计能力。它的价值不在单一 SOTA,而在于把 genomic modeling 推向真正的平台层。

局限/为何不再升一级:尽管论文层级和平台属性都很强,但是否达到 AlphaFold 那种范式重排级影响,还要看社区复现、下游采用和真实生物设计闭环的持续验证。因此当前更稳妥地定为颠覆性,而不是直接升到范式级。

原始摘要

All of life encodes information with DNA. Although tools for genome sequencing, synthesis and editing have transformed biological research, we still lack sufficient understanding of the immense complexity encoded by genomes to predict the effects of many classes of genomic changes or to intelligently compose new biological systems. Artificial intelligence models that learn information from genomic sequences across diverse organisms have increasingly advanced prediction and design capabilities . Here we introduce Evo 2, a biological foundation model trained on 9 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life to have a 1 million token context window with single-nucleotide resolution. Evo 2 learns to accurately predict the functional impacts of genetic variation—from noncoding pathogenic mutations to clinically significant BRCA1 variants—without task-specific fine-tuning. Mechanistic interpretability analyses reveal that Evo 2 learns representations associated with biological features, including exon–intron boundaries, transcription factor binding sites, protein structural elements and prophage genomic regions. The generative abilities of Evo 2 produce mitochondrial, prokaryotic and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods. Evo 2 also generates experimentally validated chromatin accessibility patterns when guided by predictive models and inference-time search. We have made Evo 2 fully open, including model parameters, training code , inference code and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity.

解读视频

相关论文

链接