Sunday, September 6, 2026

News

Google DeepMind and Wellcome Sanger Institute Launch AI Consortium for Genomics Research

ResearchPatryk Raba
Google DeepMind and Wellcome Sanger Institute Launch AI Consortium for Genomics Research
Fot. Deepa Manthravadi (DesiLady), Wikimedia Commons (CC BY-SA 3.0)

Google DeepMind and Google.org will commit $25 million over five years to a joint consortium with the Wellcome Sanger Institute, aimed at building open genomic datasets to train AI models in biology.

Contents
  1. Goals of the consortium
  2. Open access as a condition
  3. Building on earlier collaboration
  4. Implications for biomedical research

Google DeepMind, Google.org and the UK-based Wellcome Sanger Institute have announced a new consortium that will spend the next five years building high-quality genomic datasets designed for training artificial intelligence models in biology. The initiative was unveiled on June 8, 2026, at the AI x BIO conference.

Goals of the consortium

According to the announcement from both institutions, the consortium will focus on closing gaps in available biological data through strategic generation of new datasets. The aim is data designed from the ground up for training machine learning models that predict biological processes, rather than repurposing existing databases collected for other purposes.

The Wellcome Sanger Institute is one of the world's leading genome sequencing organizations, known among other things for its role in the international Human Genome Project. Google DeepMind has spent years developing biological models, including the AlphaFold family for predicting protein structures, so pairing Sanger's data infrastructure with DeepMind's modeling expertise is meant to accelerate the next generation of AI tools for the life sciences.

Open access as a condition

The institutions publicly stated that all data and methodological frameworks developed through the consortium will be openly available, without the licensing restrictions typical of commercial databases. That sets the initiative apart from many earlier partnerships between big tech and research centers, in which data remained closed or accessible only under contractual terms.

By combining Sanger's expertise in building datasets with Google DeepMind's edge in AI, we can accelerate the generation of biological data for foundational AI models - Dr. Julia Wilson, Chief Innovation and Impact Officer, Wellcome Sanger Institute
Together we want to build the data backbone needed to decode the complexity of biological processes - Dr. Pushmeet Kohli, Vice President of AI for Science, Google DeepMind

Building on earlier collaboration

The new consortium is not starting from scratch. Sanger and Google DeepMind had already worked together on, among other things, a fellowship program focused on AI applications in genomics and efforts to strengthen AI capacity in low- and middle-income countries. The current partnership formalizes and scales those efforts, giving them a five-year funding horizon.

The philanthropic side is handled by Google.org, the company's division that funds social and scientific projects. Its representative stressed that the priority is building open data foundations that future generations of biological models can draw on, not necessarily only those built by Google.

We're supporting open data foundations that will power the next generation of biological AI models - Leslie Yeh, Director of Science Advancement, Google.org

Implications for biomedical research

The lack of well-labeled, high-quality biological datasets has long been considered one of the main barriers to developing AI models for the life sciences, unlike text or images, where far more raw training data is available. The consortium aims to address exactly this problem by generating data where it has been missing, rather than solely processing existing archives.

The effects of such programs can take years to materialize, but open data access lowers the barrier to entry for smaller research teams, including Polish academic and biotech centers, which typically lack budgets comparable to large pharmaceutical labs. Open genomic datasets could therefore indirectly support domestic projects in computational biology as well.

The consortium said it will accept additional institutional partners as the program moves forward, suggesting the initiative's final scope and budget could still grow in the coming years.

Share: