Research2026-08-27

Google DeepMind ran what it describes as the first double-blind test of a commercial frontier AI model, using encrypted computing hardware so the outside testers cannot see the model's internals and Google cannot see their test questions. The pilot used a Gemini Flash Lite model with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. It remains a pilot rather than a standing offer, and a technical report accompanies it.

What changed

External testers either received the test prompts risk of the model owner seeing them, or the model owner had to hand over model weights.

What it unlocks

Independent organizations can test a commercial model against secret benchmarks without revealing either the questions or the model's internals.

What you need to act on it

  • participation as a partner organization in the pilot

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.