Google DeepMind ran what it describes as the first double-blind test of a commercial frontier AI model, using encrypted computing hardware so the outside testers cannot see the model's internals and Google cannot see their test questions. The pilot used a Gemini Flash Lite model with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. It remains a pilot rather than a standing offer, and a technical report accompanies it.
What changed
External testers either received the test prompts risk of the model owner seeing them, or the model owner had to hand over model weights.
What it unlocks
Independent organizations can test a commercial model against secret benchmarks without revealing either the questions or the model's internals.
What you need to act on it
- participation as a partner organization in the pilot
Sources