Saturday, July 25, 2026

News

DeepMind Warns: AI Agents Are Outpacing Science's Ability to Verify Discoveries

ResearchPatryk Raba
DeepMind Warns: AI Agents Are Outpacing Science's Ability to Verify Discoveries
Fot. Gciriani, Wikimedia Commons (CC BY-SA 4.0)

A new Google DeepMind policy report finds that AI agents are generating scientific hypotheses faster than labs can check them, and calls for urgent investment in validation infrastructure.

Contents
  1. What DeepMind's Systems Can Do
  2. The Verification Problem
  3. Recommendations for Policymakers and Institutions
  4. What It Means for Science and Poland

Google DeepMind has published a policy report describing a new problem emerging as AI agents become more common in science: these systems can generate research hypotheses and propose experiments faster than laboratories can verify them. The authors call this phenomenon the validation bottleneck and call for urgent changes in how scientific infrastructure is funded.

The report was authored by Don Wallace, Conor Griffin, Sean O'Neill, Thang Luong, and Owen Larter of DeepMind's public policy team. Its starting point is the observation that an AI agent differs from an ordinary chatbot in that it can independently plan a path to a goal: breaking a task into stages, running multiple subprocesses in parallel, and correcting its own mistakes along the way. Applied to science, this means systems that combine hypothesis generation, experiment design, and data interpretation on a single platform.

What DeepMind's Systems Can Do

The authors describe several tools already in use. Co-Scientist, a multi-agent system built on Gemini, iteratively generates, debates, and evolves new research hypotheses. It has been demonstrated in three biomedical areas: drug repurposing for acute myeloid leukemia, the search for new therapeutic targets in liver cirrhosis, and explaining mechanisms of antibiotic resistance.

AlphaEvolve, another system mentioned in the report, orchestrates teams of agents that optimize algorithmic solutions across vast solution spaces - it has supported work including TPU chip design and genomic analysis. The authors also cite the case of mathematician Terence Tao, who used AlphaEvolve's algorithmic optimization in his own research.

The Verification Problem

The report's central claim is that the physical and institutional limits on experimental verification have not sped up to match the pace at which AI generates ideas. Laboratories, peer reviewers, and research infrastructure still operate at pre-agent speed, creating a jam between the volume of proposals and science's ability to confirm them.

Researchers can suddenly do things they simply couldn't do before - Matej Balog
One fabricated claim on page ten of a result can invalidate the whole thing - Vivek Natarajan

Aletheia, a system that pairs mathematical proof generators with natural-language verifiers, solved 6 of 10 unpublished problems in the First Proof challenge. Thang Luong, one of the report's authors, warns of what he calls 'proof indigestion' - a situation where the number of generated solutions exceeds humans' capacity to check them.

Recommendations for Policymakers and Institutions

The report sets out four main recommendations. First, ensuring broad access to AI agents through public-private partnerships, modeled on earlier supercomputer access programs. Second, preparing data infrastructure - making open datasets available through documented APIs with quality control and metadata standards.

Third, the authors call for investment in experimental validation infrastructure, including automated laboratories and expanded public facilities - citing as an example the $100 million from the US NSF and the £81 million invested in the UK's Materials Innovation Factory mentioned above. Fourth, they propose reworking the scientific peer review system: letting reviewers use agents themselves while introducing transparency requirements, including watermarking and mandatory disclosure of AI use.

What It Means for Science and Poland

For Polish universities and research institutes, the DeepMind report signals that access to agentic tools and validation infrastructure could soon become a factor in international scientific competitiveness. The authors stress that the current moment is a 'formative window' requiring serious structural investment in science's institutional foundations, before an edge in generating hypotheses turns into a chaos of unverifiable results.

The report is not the announcement of a new product but a voice in the policy debate over how grant agencies, scientific journals, and universities should respond to the pace that systems like Co-Scientist and AlphaEvolve are bringing to research. Sources: Conjecture Machines: AI agents and the new validation bottleneck in science (deepmind.google).

Share: