Ad
Skip to content

AI benchmarks have a trust problem and Google wants to fix it

Image description
GPT-Image-2 prompted by THE DECODER

Google Deepmind wants to use a cryptographic method to stop AI models from seeing test questions in advance. A pilot project with the Singapore AI Safety Institute and other partners runs a double-blind test on a Gemini model.

If a test-taker knows the questions ahead of time, even a perfect score is worthless. Google Deepmind uses this image to describe a core problem in evaluating AI models: benchmark contamination. If a model has already seen the test questions during training, you can only trust the results so far.

To prevent that, Google Deepmind says it's launching the first double-blind evaluation of a proprietary frontier AI model. External tests stay locked in a cryptographic "box," so a model can't later use those questions to optimize itself specifically for the test.

For the pilot, Google is testing a model from the Gemini Flash Lite line against confidential benchmarks.

The tradeoff the method aims to fix

Highly sensitive external evaluations used to require a compromise, Deepmind says. Either the evaluators handed over their test prompts, which let the model provider see the questions in advance. Or the provider handed over its model weights and risked its intellectual property. A recent example of this dilemma was the delayed evaluation for the ARC-AGI benchmark of Anthropic's Fable 5, since the AI company enforces a 30-day data retention policy for its strongest models.

The double-blind evaluation is meant to eliminate that tradeoff. Google uses Confidential Space from Google Cloud's confidential computing portfolio to do it. The setup cryptographically verifies that both the external test data and the model stay private to their respective owners. The evaluator never sees the Gemini weights, and Google never sees the test prompts.

Until now, zero-logging protocols and contractual safeguards kept external prompts confidential. Adding technical and cryptographic protection is a big step forward for secure model evaluation, the company says.

Why this matters most for sensitive areas

The cryptographic proof aims to prevent contamination and protect sensitive data. Deepmind says this matters most for highly sensitive evaluations, such as cybersecurity or tests run by government agencies. Independent organizations could rigorously test advanced models without giving up data sovereignty or security.

Google hopes the effort sets a new standard for model oversight and helps the industry build more reliable and widely trusted AI systems. Google lays out the details on methodology and results in a technical report.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: "AI Radar" — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI
Subscribe to The Decoder