Saturday, September 12, 2026

News

Google DeepMind: AI Agents Generate Scientific Hypotheses Faster Than They Can Be Verified

ResearchPatryk Raba

In a new science-policy analysis, Google DeepMind describes how AI agents are beginning to reshape scientific research much as they earlier reshaped software development, and warns of a widening gap between the speed of hypothesis generation and science's capacity to verify it.

Contents
  1. A decade of work, verified in two days
  2. More examples from medicine and mathematics
  3. The validation bottleneck
  4. Recommendations for science policy

In July 2026, Google DeepMind published an analysis titled "Conjecture Machines: AI Agents and the New Validation Bottleneck in Science," in which five of the company's authors describe how systems like Co-Scientist are starting to change the way scientists work, and why the ability to generate new hypotheses is growing far faster than laboratories' capacity to verify them.

A decade of work, verified in two days

The analysis's central example is the story of microbiologist José Penadés of Imperial College London. His team spent nearly ten years determining how a certain family of superbugs transfers antibiotic resistance between species, a finding that remained unpublished and known only within his lab.

In 2024, Penadés described the problem to Co-Scientist, a multi-agent AI tool built by Google DeepMind on top of Gemini models. Within two days, the system produced five ranked hypotheses, and the top-ranked one matched the team's unpublished conclusion: that some superbugs acquire viral "tails" that let them jump between host species. According to the analysis, Penadés was surprised enough by the match that he contacted Google to check whether the tool had accessed his unpublished data; the company confirmed no leak had occurred.

More examples from medicine and mathematics

The analysis also describes the case of Gary Peltz of Stanford, who used Co-Scientist to search for drugs that could be repurposed to treat liver fibrosis. His own two proposed candidates showed no effect in lab tests, while two of the three candidates identified by the system halted fibrosis progression and supported liver regeneration.

Other tools mentioned include AlphaEvolve, used among other things for chip design, for supporting mathematician Terence Tao's work on open Erdős problems, and for genomic analysis. Aletheia, a system dedicated to mathematical proofs, solved six of ten previously unpublished problems during the inaugural First Proof challenge.

The validation bottleneck

The document's central argument is that AI agents are now far more effective at generating ideas and candidate solutions than at verifying them, a process that still requires physical experiments, time, and reviewers' work. The authors call this the "new validation bottleneck," a growing gap between the pace at which new hypotheses are produced and the pace at which science can check them.

A single fabricated detail buried deep in a result can undermine the entire finding - Vivek Natarajan, Google DeepMind
Research will shift toward high-level orchestration rather than manual lab work - Natasha Latysheva, Google DeepMind

Recommendations for science policy

The document sets out four recommendations for policymakers. The first is to guarantee broad access to AI agents for scientists, treated as a strategic priority comparable to the historical push to provide access to supercomputers, citing the US Genesis Mission initiative as a possible funding model.

The second recommendation concerns preparing scientific data for use by agents, including building accessible APIs for open datasets and privacy-protecting frameworks, citing the OpenSAFELY project as an example for sensitive data such as genomic data. The third recommendation is to expand experimental infrastructure and validation capacity, including automated laboratories; the authors point to Google DeepMind's lab partnership with the Francis Crick Institute and funding examples from the NSF and UK institutions.

The fourth recommendation concerns reforming scientific peer review: equipping reviewers with AI tools and developing standards for disclosing AI use, such as watermarks or so-called Human-AI Interaction Cards, while keeping human judgment central to evaluating publications.

Thang Luong, one of the analysis's authors, warns in the document of a phenomenon he calls "proof indigestion," a situation in which the pace at which AI generates results outstrips human reviewers' capacity to process them. Alex Davies, meanwhile, describes a possible scenario in which machines take over part of the discovery process while humans focus on deciding which research directions are worth pursuing further.

For Poland's universities and research institutions, which have spent months grappling with questions about AI's role in teaching and publishing, the DeepMind document makes the point plainly: the key constraint will no longer be access to the models themselves, but institutions' ability to verify what those models produce. That is a question of lab infrastructure, review funding, and standards for disclosing AI's role in research, not just of access to tools.

Share: