Friday, August 28, 2026

News

Google DeepMind's Co-Scientist Now Plans Experiments and Writes Research Papers

ResearchPatryk Raba
Google DeepMind's Co-Scientist Now Plans Experiments and Writes Research Papers
Fot. Gciriani, Wikimedia Commons (CC BY-SA 4.0)

Google DeepMind has expanded its Co-Scientist system from a hypothesis generator into a full lab partner that plans experiments, operates equipment, and writes its own manuscripts. Tests across three fields cut research timelines from days to minutes, but also revealed the system's tendency toward scientific confabulation.

Contents
  1. What Changed
  2. Three Fields, Three Levels of Autonomy
  3. The Confabulation Problem
  4. Context for Science and Business

Google DeepMind announced on August 28, 2026 an expansion of its Co-Scientist system from a tool that generates research hypotheses into a full, closed-loop research pipeline. The new version of the system plans experiments, operates laboratory equipment, writes executable code, analyzes results, and independently produces scientific manuscripts.

What Changed

Co-Scientist debuted in May 2026 as a system for generating and ranking research hypotheses, described in a paper published in Nature. At the time, the tool didn't run experiments itself, it only flagged which ones were worth testing, drawing on literature reviews and structured databases such as ChEMBL and UniProt. The August update closes that loop: a system that once only proposed ideas now plans experimental protocols itself, controls lab instruments, and writes the final text of the research paper.

Co-Scientist's architecture relies on several specialized agents working across three phases: generating hypotheses grounded in the literature, critically evaluating them through "idea tournaments" scored in a manner similar to Elo rankings, and further refining the top-rated proposals. According to DeepMind, most of the system's computing power goes not into generating hypotheses but into verifying them.

Three Fields, Three Levels of Autonomy

In materials science, Co-Scientist designed synthesis recipes for two-dimensional materials. After 25 rounds of human-assisted refinement, the system correctly synthesized thin semiconductor films on its first attempt, which DeepMind says cut recipe development time from days to minutes.

In biology, the system independently built an image-analysis pipeline predicting E. coli bacterial colony patterns, matching previously unpublished results on three of four measured shape traits. The furthest-reaching autonomy came in computer science: there, Co-Scientist operated entirely on its own and designed the architecture of Agent_H, a medical AI system that outperformed six leading competing models on health-related benchmarks.

The Confabulation Problem

A key part of the update is a set of reliability modules designed to curb result fabrication, a common problem with language models used to generate scientific content. The verification mechanism cross-checks numerical claims against actual code execution logs. With the modules enabled, the system fabricated key results in 4 percent of cases, versus 46 percent without them and as much as 90 percent in comparable competing systems.

Despite the improvement, DeepMind researchers acknowledge the problem hasn't disappeared entirely. The system shows a tendency toward selective reporting of results and, as noted in the documentation, "writes highly plausible-sounding methods sections that didn't match its actual code." Human evaluation of Agent_H's medical responses found statistical significance in only one category, harm reduction, despite strong performance on standard benchmark tests.

Context for Science and Business

The earlier, May version of Co-Scientist reached individual researchers through the Gemini for Science tool, while companies such as Daiichi Sankyo, Bayer Crop Science, and US national laboratories tested enterprise versions. Even then, the system had six published papers to its name: a Stanford team used it to find drug-repositioning candidates for liver fibrosis, one of which blocked 91 percent of scarring reactions; MIT and Harvard researchers developed RNA-based approaches to treating ALS; and a Cambridge team narrowed the search for proteins linked to zoonotic diseases.

Co-Scientist is entering a market where competitors such as Lila Sciences, Medra.ai, Phylo's Biomni Lab, and AWS BioDiscovery are already active. Until now, DeepMind's system stood out for its focus on the idea-generation phase rather than autonomously running lab experiments. The August update blurs that distinction, pushing the tool toward full automation of the research cycle.

Most of the system's computing power is dedicated to verifying hypotheses, not generating them - Google DeepMind, Co-Scientist documentation

For Polish research institutions and biotech companies, tools like this signal a potential reduction in the time between a research idea and its laboratory validation, but also the need to develop quality-control procedures for AI-generated results. The persistent, if significantly reduced, rate of fabricated results shows that human oversight of the system's conclusions remains essential, especially in medical applications.

DeepMind has not yet given a timeline for a broader rollout of the expanded, fully automated version of Co-Scientist beyond the testing environment. The company notes that fully closing the research loop requires integration with lab automation platforms, which for now limits deployment to partners with the right equipment.

Share: