Google DeepMind unveils pilot for double-blind AI evaluation
Google DeepMind has introduced a pilot project where AI models are evaluated within a cryptographically isolated environment to prevent data leakage and protect model weights.
What happened?
Google DeepMind has presented the results of a pilot project for the double-blind evaluation of AI models. By utilising a Trusted Execution Environment (TEE) based on Google Cloud Confidential Space and Nvidia H100 hardware, developers and external auditors were isolated from one another. In the pilot, Google provided model weights for Gemini Flash Lite, while external organisations provided confidential test data. This method ensures that evaluators cannot access the model weights, while the model developer cannot see the test queries.
Key facts
Why it matters
Leakage of test data into training sets, known as benchmark contamination, is a well-recognised issue in AI research that complicates the measurement of a model's true capabilities. Simultaneously, model developers often seek to avoid disclosing sensitive source code or model weights to external parties. The double-blind method offers a technical solution to this conflict by allowing evaluation to occur within a cryptographically secure environment.
Who is affected?
The method concerns AI developers, independent safety institutes, and evaluation organisations. Participants in the pilot project included the Singapore AI Safety Institute, MLCommons, OpenMined, and AVERI. The system is primarily relevant to entities that develop or evaluate advanced AI models where both intellectual property and benchmark integrity must be safeguarded.
Impact on the EU
The pilot project is based on cloud infrastructure that can be deployed in data centres globally, including within the EU. How such solutions align with the requirements for auditing and transparency under the EU AI Act remains to be seen, as the regulation imposes strict transparency requirements for AI models deemed to carry systemic risk.
What else you should know
The pilot project was conducted as a collaboration to practically test how cryptographically isolated evaluations can function in large-scale environments. Future tests will determine if the method can scale to larger model architectures and more complex testing protocols.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka organisationer deltog i piloten?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Google DeepMind unveils pilot for double-blind AI evaluation"