Friday, August 28, 2026

News

Google DeepMind Tests AI Model Using Double-Blind Trial

ResearchPatryk Raba
Google DeepMind Tests AI Model Using Double-Blind Trial
Fot. Gciriani, Wikimedia Commons (CC BY-SA 4.0)

Google DeepMind has run the world's first double-blind evaluation of a frontier-class AI model, in which neither the evaluator nor the model's creator can see the other side's data. The pilot, involving Gemini Flash Lite, aims to solve the problem of benchmark contamination.

Contents
  1. The Benchmark Problem
  2. How the Cryptographic Box Works
  3. Pilot Partners
  4. Why It Matters for AI Evaluation

Google DeepMind announced on August 27, 2026 a pilot of a new method for testing artificial intelligence models, in which neither the external evaluator nor the model's creator has access to the other party's data. It is the world's first double-blind evaluation of a frontier-class model, conducted on Gemini Flash Lite.

The Benchmark Problem

AI test results only mean something if the model has never seen the exam questions before. As models grow more capable, there is a growing risk that benchmark content leaks into training data, whether deliberately or by accident. A high score then stops measuring the model's actual abilities and starts measuring simple memorization.

DeepMind compares this to a student who knows the exam questions in advance. Even if he answers correctly, the score says nothing reliable about his actual knowledge. Until now, solving this problem required a trade-off: either the AI company had to disclose the model's weights to an independent evaluator, or the evaluator had to hand over its test questions to the company, which ruined their confidentiality for future use anyway.

How the Cryptographic Box Works

The new method relies on Confidential Space technology, part of Google Cloud's Confidential Computing portfolio. It creates an isolated, encrypted computing environment where the model and the test set meet without revealing their contents to any external party, including Google itself.

In practice, this means the evaluator has no access to Gemini Flash Lite's model weights, and Google cannot see the questions prepared by the external institution. Once testing is complete, both sides receive cryptographic proof that their data remained private throughout the process, without having to rely on mutual trust based solely on declarations.

Double-blind evaluations eliminate this compromise - Google DeepMind

Pilot Partners

Four organizations working on AI safety and evaluation are taking part in the project: the Singapore AI Safety Institute, the nonprofit OpenMined, which specializes in privacy-preserving technologies, AVERI, and MLCommons, the consortium known for running the industry-standard MLPerf benchmarks. Each brings a different perspective to testing the integrity of results.

The choice of Gemini Flash Lite as the first model tested is not accidental. It is a smaller, cheaper-to-run model from the Gemini family, which lets DeepMind test the entire technical process at lower risk before applying the method to flagship frontier-class models.

Why It Matters for AI Evaluation

Benchmark contamination is a problem the entire industry has grappled with for years. Regulators, researchers, and companies using AI models must rely on published test results to assess a given system's real capabilities and safety. If those results can be artificially inflated, the whole mechanism for comparing models loses credibility.

Double-blind evaluation addresses the problem at its source, eliminating the possibility that a producer's model encounters test content before the official measurement, while also protecting the uniqueness of the evaluator's questions for future use. This also matters for independent AI safety institutes, which need to maintain the credibility of their tests across multiple model producers at once.

For Polish companies and institutions using Google's models, for example through cloud services or Workspace integrations, more credible benchmarks mean a better basis for choosing an AI vendor. If the method proves effective and is extended to other models and other producers, it could become a new industry standard for assessing the capabilities and safety of AI systems, replacing the more easily disputed rankings used until now.

DeepMind has not yet given a specific timeline for extending the program to further models or for publishing the results of the Gemini Flash Lite pilot. The company has only said it views this as a first step toward evaluation infrastructure that is more resistant to manipulation for the entire industry.

Share: