We've updated our Privacy Policy to make it clearer how we use your personal data. We use cookies to provide you with a better experience. You can read our Cookie Policy here.

Advertisement

The Trillion Gene Atlas: Unlocking Earth's Hidden Genetic Diversity

Black and white illustration of a DNA double helix unwinding and forming a binary code.
Credit: iStock.
Read time: 5 minutes

Artificial intelligence (AI) is reshaping how scientists discover and develop new therapeutics, but the performance of biological AI models depends on the quality and diversity of the data on which they are trained. Many current models rely on relatively limited genomic datasets that capture only a small fraction of Earth's biological diversity.


To address this challenge, Basecamp Research has launched the Trillion Gene Atlas, an ambitious initiative that aims to expand known evolutionary genetic diversity by 100-fold through the collection and analysis of genomic data from more than 100 million species worldwide.

 

Developed in collaboration with PacBio, Ultima Genomics, Anthropic, and NVIDIA, the project intends to create a vast new resource for training next-generation biological AI models.


Technology Networks recently spoke with Dr. Glen Gowers, co-founder and CEO of Basecamp Research, and Christian Henry, CEO of PacBio, to learn more about the Trillion Gene Atlas and its significance.

 

In this interview, they also discuss the importance of long-read sequencing for improving data quality, the technical and computational challenges associated with building datasets at scale, and how the project could transform our understanding of evolution, biodiversity, and therapeutic discovery.

Anna MacDonald (AM):

What are the main goals of the Trillion Gene Atlas, and how does mapping a trillion genes across diverse ecosystems provide new opportunities for understanding evolution, gene function, and biodiversity?


Glen Gowers, PhD (GG):

We hope to prove that with enough data points, it’s possible to push biological AI models towards a new territory of performance. Our team already showed the value of scale with our EDEN foundation model in January 2026, which was trained on 10 trillion nucleotides of novel evolutionary data.

 

Since this represented about an order of magnitude more data than is publicly available, we were able to establish new scaling laws that show how generalized AI model “intelligence” scales with dataset size.

 

The Trillion Gene Atlas continues to build on this work, representing about three orders of magnitude, or 1,000 times, more training data than is publicly available. We expect scale in datapoints to be reflected in scale of novel biological insights we can unlock. 



Christian Henry (CH):

With the Trillion Gene Atlas, PacBio aims to prove that long-read sequencing can scale to industrial levels and that the technology is no longer confined to specialist use cases.

 

Our lab will deliver complete, information-rich genomes on a scale even larger than the ambitious Darwin Tree of Life Project, to support researchers’ understanding of evolutionary relationships, gene function, and previously uncharacterized biology.

 

The project’s results should demonstrate that the accuracy and granular detail provided by long-reads are non-negotiable when building AI models.



AM:
A lot of today’s biological models are trained on limited or biased public datasets. What limitations does this create, and how will the Trillion Gene Atlas address these issues?

GG:

Today’s genomic datasets are skewed toward a limited set of organisms from well-explored environments, with almost 70% of all public data coming from only five species. But there are plenty of remote landscapes rich in plant and microbial life that remain largely uncharted, despite holding much of nature’s functional diversity.

 

As a result of limited training data, most of today’s AI models don’t truly understand biology yet. It’s the equivalent of training ChatGPT on only a few books. The model is good at recognizing familiar patterns but struggles the moment you ask it to hypothesize something new, such as a novel protein, a rare organism, or a different environment.

 

That’s why at Basecamp, our efforts are twofold. Firstly, working with a network of partners and field teams to expand the breadth of data we have on species. Then we’ll use this data to build an AI trained on a scale like never before, including billions of previously unknown genes across more than a million new-to-science species. 



CH:

Beyond the diversity issue that Glen raises, genomic datasets are often limited by data quality and depth of genomic insight. Many species’ genomes remain incomplete, with complex regions, tandem repeats, and highly homologous genes missed or mischaracterized by traditional sequencing approaches. This means AI models are often trained on fragmented representations of biology, increasing the risk of biased or inaccurate insights.

 

The Trillion Gene Atlas addresses this challenge by using PacBio HiFi long-reads to generate comprehensive and accurate genomes with the level of resolution needed to train reliable AI models. 



AM:
The project is enabled by advanced sequencing technologies. What advantages does long-read sequencing provide over short-read approaches? 

CH:

Long-read sequencing preserves long stretches of native DNA, enabling accurate assembly of complete genomes, including complex and repetitive regions. By preserving genomic context, HiFi reads deliver both more complete and accurate views of species’ underlying biology, down to the subspecies or strain level.

 

Compared to short-reads, PacBio’s approach ensures the data reflects the true complexity of biology rather than a fragmented approximation, which is critical for downstream applications such as AI model training and drug discovery. 



AM:
What are the main biological, technical, and computational bottlenecks you anticipate in this project, and how are the project partners working to overcome them?

GG:

From a biological standpoint, accessing and preserving diverse, high-quality samples is challenging in complex environments. Our field scientists are highly experienced in collecting DNA samples in a standardized way in multiple extreme environments—from ice caps to Costa Rican rainforests.

 

On the technical side, analyzing this scale of genomic data simply isn’t possible manually. AI is central to enabling scientists to explore this data and learn the underlying rules of biology rather than simply catalog it.


Our approach is to create AI that understands how evolution solves problems, to help develop new therapeutic solutions from a disease prompt. Instead of screening existing compounds, the model can generate novel proteins, peptides, or generic designs informed by evolutionary datasets.

 

Finally, AI on this scale requires significant compute power. To scale the compute and reasoning required to support the Trillion Genome Atlas and make it practical, Basecamp has partnered with Anthropic and NVIDIA. 



CH:

Increasing throughput while maintaining accuracy and reducing cost per genome remains a common technical challenge. At PacBio, we’ve made significant progress in moving long-reads from a specialist tool to operating at scale, and we’ve also increased capacity at our HQ lab in Menlo Park to support Basecamp’s project.

 

Improvements across the broader workflow, from sample processing to computing, have eased analysis bottlenecks, making it feasible to generate the volume and quality of data required for initiatives like the Trillion Gene Atlas. 



AM:
How are Digital Sequence Information regulations shaping decisions around data governance, equitable access, and international collaboration within the Trillion Gene Atlas?

GG:

Digital Sequence Information regulations are shaping the Trillion Gene Atlas by guiding fair data governance, equitable access, and global collaboration. We’ve built a network of scientific collaborators across 31 countries, grounded in local capacity building, knowledge exchange, and equitable Access and Benefit-Sharing agreements. This framework ensures high-quality genomic data collection, continued investment in scientific infrastructure, and training within partner regions.

 

These partnerships reflect that both the opportunity and responsibility of building a resource like the Trillion Gene Atlas are far greater than any one organization, requiring global coordination and collaboration to ensure the project benefits both science and society.



CH:

Ensuring equitable access to high-quality genomic data is essential to maximize its scientific and societal impact. More than 90% of publicly available genome-wide association studies (GWAS) data is derived from participants of European descent, and 86% of large genomic studies include only one ancestry group.

 

This approach severely limits the applicability of results, underlining the need for more equitable research across the genomics industry. Efforts such as PacBio’s support in building the first Arab human pangenome and for the Korean Pangenome Reference Project are positive steps in this direction, and we look forward to continuing this work with partners.



AM:
What kinds of scientific or therapeutic challenges do you think this project could help to tackle? Are there specific areas where you anticipate the greatest impact?

GG:

The impact will be broad because this dataset expands the biological space AI can learn from by orders of magnitude. Today, there are tens of thousands of genetic diseases for which we know exactly which gene is involved, but we still lack the ability to design effective treatments. By capturing far greater genetic diversity, we can move from simply predicting what biology does to designing new biological solutions with intention and high predictability.

 

Early breakthroughs will likely come in drug discovery, where richer data can unlock new targets and therapeutic approaches. Over time, the same approach will apply far more widely, from industrial biotech to environmental applications.



CH:

We expect the first breakthroughs to happen in areas where complex biology has limited progress, such as rare disease and oncology. By providing AI models with accurate, high-resolution data, there will be new opportunities to explore previously uncharacterized genes and pathways across diverse species. These new biological insights can accelerate target discovery and support more confident therapeutic design.  


The introduction to this interview includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.



Google News Preferred Source Add Technology Networks as a preferred Google source to see more of our trusted coverage.