Benchmark contamination — when evaluation prompts or questions become accessible to models before testing — undermines the reliability of model assessments. To address this, Google and several external partners launched a pilot they describe as the world’s first double-blind external evaluation of a proprietary, frontier-class AI model. The model under test is Gemini Flash Lite.
Participants
External partners in the project include Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. Google says the initiative adds cryptographic safeguards to protect both proprietary models and confidential evaluation data, with the aim of increasing the integrity of external assessments.
How the approach differs
Traditionally, external high-stakes evaluations required a trade-off: evaluators would hand over their test prompts (risking that the model provider could see them), or the model provider would share model weights (risking disclosure of intellectual property). The double-blind setup removes that compromise. By running both the external evaluation data and the proprietary model inside Confidential Space from Google Cloud’s Confidential Computing portfolio, cryptographic techniques are used to verify that each party’s data remains private to its owner.
In this arrangement, the evaluator cannot access the Gemini model weights, and Google cannot view the evaluator’s test prompts. Google frames these technical and cryptographic protections as a step beyond prior confidentiality approaches that relied primarily on non-disclosure protocols and contractual safeguards.
Why this matters
The method reduces the risk of benchmark contamination and protects sensitive evaluation content, which becomes increasingly important as models grow more capable. This is particularly relevant for high-sensitivity evaluations, such as those used in cybersecurity contexts or by government bodies. Cryptographic assurance enables independent organizations to carry out rigorous tests without compromising data sovereignty or model confidentiality.
Goals and next steps
Google expects this pilot to help set a new standard for model oversight, supporting the development of safer and more widely trusted AI systems. Further details on methodology and results are provided in a technical report published by Google.



