Google DeepMind urges funders to tackle AI science validation gap

Google DeepMind sets out four policy priorities as AI agents accelerate scientific ideation, creating new pressures for research infrastructure, graduate training and peer review.

Abstract AI network of molecules and data particles narrowing through a glass bottleneck toward a laboratory petri dish

Google DeepMind says AI agents can generate scientific hypotheses faster than researchers can validate them through physical experiments

Google DeepMind has published a four-part policy agenda calling on governments and science funders to widen access to AI agents, prepare national datasets for agent use, expand experimental infrastructure and equip peer reviewers with AI tools.

The proposals respond to what the company describes as a growing “validation bottleneck” in science. Agentic AI can generate hypotheses and candidate solutions at scale, but testing those ideas in laboratories and assessing them through existing research processes remains slow and expensive.

The article, Conjecture Machines: AI agents and the new validation bottleneck in science, is written by Don Wallace, Conor Griffin, Sean O’Neill, Thang Luong and Owen Larter. It draws on discussions with 10 Google DeepMind researchers and engineers.

Sharing the piece on LinkedIn, Luong, Principal Scientist and Director of Research at Google DeepMind, wrote: “Ideas are becoming abundant, but testing them in the physical world remains slow and costly. For policymakers, funders, and researchers to truly unlock superhuman scientific discovery, we must understand and address this growing gap between generation and verification.”

He identified four priorities for policymakers and funders: widening access to AI agents, making national data assets agent-ready, tackling the validation bottleneck and equipping peer reviewers with agents.

The central argument is that AI agents are changing the balance between generating scientific ideas and verifying them. The authors describe the systems as “conjecture machines” capable of reviewing literature, producing hypotheses, writing code and coordinating multiple processes, while physical experiments and institutional review continue to set the pace.

AI-generated ideas still require human verification

Google DeepMind cites its Co-Scientist agent as an example of the speed difference. Microbiologist José Penadés and his team at Imperial College London spent most of a decade investigating how a family of superbugs spreads antibiotic resistance. According to the company, Co-Scientist produced five possible explanations within two days after Penadés described the problem, with its highest-ranked hypothesis matching the conclusion reached by his laboratory.

The article also describes Stanford University researcher Gary Peltz’s use of Co-Scientist to identify existing drugs that could potentially be repurposed for liver fibrosis. Peltz selected two candidates and the agent selected three. Neither of his choices showed a benefit in assays using live human liver cells, while two of the agent’s candidates blocked fibrosis and promoted liver cell regeneration, according to the account.

These examples come from a Google DeepMind article rather than an independent evaluation of the systems. The authors also acknowledge that agents remain unreliable and that detecting an incorrect claim can require substantial subject expertise.

“A single hallucinated claim on page 10 of an output can invalidate the whole thing,” says Vivek Natarajan, a Co-Scientist lead.

The company argues that scientific agents should expose their reasoning, communicate uncertainty and provide evidence for their conclusions rather than operate as systems that only produce answers. Calibrating confidence across open-ended scientific reasoning remains unresolved.

Mathematics and computer science offer more opportunities for automated checking because outputs can sometimes be verified through code or formal languages. At the inaugural First Proof challenge in February, Google DeepMind’s Aletheia system solved six of 10 unpublished research problems within one week, which the company describes as the strongest result.

That creates a different capacity problem for mathematicians.

“We are moving toward a future of serious 'proof indigestion' where AI generates breakthroughs faster than humans can review them,” says Luong, who led the Aletheia work.

Funding access, data and experimental infrastructure

The first policy recommendation concerns access. Google DeepMind argues that science funders should determine how laboratories can select and pay for agents, including whether existing research grants are sufficient or new funding programs are required.

The authors compare the issue with providing researchers access to supercomputers and call for public-private partnerships capable of supplying AI systems at the required scale. The article does not provide projected costs or an implementation timetable.

Its second recommendation is to make publicly funded data easier for agents to use. Open or lower-risk datasets should be accessible through documented APIs, supported by metadata, quality controls and interoperable standards, according to the authors.

More sensitive information, including genomics and virology data, would require privacy controls, auditability and restrictions governing how agents access it. The article points to OpenSAFELY, which provides researchers with controlled access to health data, as one possible model.

The third priority is increasing the capacity to test AI-generated ideas. Google DeepMind calls for investment in existing public facilities alongside automated laboratories that can conduct some experiments using robotics and computing infrastructure.

The company has established a wet laboratory inside the UK’s Francis Crick Institute and says it is providing independent scientists with funding and Co-Scientist access to validate agent-generated hypotheses.

The article also cites the US National Science Foundation’s $100 million investment in a network of distributed facilities and the UK’s £81 million Materials Innovation Factory. Google DeepMind argues that the cost of automated laboratories will require governments to consider centralized facilities that researchers can access without individual institutions funding the full infrastructure.

Scientific education would also need to adjust. The authors warn that junior researchers could be concerned about institutions directing budgets toward AI usage rather than staff, while unmanaged adoption could prevent new scientists from developing the judgment required to assess agent outputs.

They suggest that science graduate training may need structured periods of work without agents, alongside access to systems that support researchers’ thinking rather than operate as unquestioned sources of answers.

AI tools proposed for peer review

Google DeepMind’s final recommendation addresses peer review. The authors argue that increasing AI use in grant applications and academic papers is making it harder for funders and reviewers to identify original thinking, verify findings and decide which work warrants support.

The article notes that some organizations are changing their processes. The UK Medical Research Council has reinstated interviews for shortlisted applicants, although the authors acknowledge that greater reliance on interviews or track records could favor established researchers and increase costs.

They propose a layered response that combines clearer disclosure of AI use with access to agents for reviewers. Suggested measures include watermarking, records known as Human-AI Interaction Cards and initial use of review agents for more objective tasks such as detecting errors.

Under the proposal, journals, funders and conferences would also require AI systems to show their reasoning, support claims with citations and document uncertainty where possible.

Previous
Previous

The Pokémon Company seeks AI agents for TCG competition on Kaggle

Next
Next

Purdue Global appoints John Higgins as Chief Transformation Officer