Every enterprise deploying AI wants confirmation that the system runs reliably and safely under its specific use cases and data. But simply connecting and running tests is far from straightforward.
Privacy and Integrity: A Dual Challenge
Deployers such as banks must protect sensitive data, while frontier labs strictly guard model weights. More critically, benchmark providers must have both technical and legal measures in place to prevent models from contaminating benchmarks through learning. Confidentiality of the evaluation alone is not sufficient; robust benchmark management procedures are also essential.
First Proof of Concept for Double-Blind Evaluation
To demonstrate top-tier benchmark integrity, MLCommons partnered with Google DeepMind, OpenMined, and AVERI to conduct the first double-blind evaluation of closed-source models. This approach leverages cryptographic safeguards to ensure evaluation components remain fully usable without contaminating the benchmark.
A held-out subset of the AILuminate™ safety benchmark was used for this validation, ensuring the Google DeepMind model had never previously encountered the evaluation set. AVERI leveraged OpenMined's secure computation to run prompts in a containerized environment, providing provable protection for both model weights and the evaluation itself.
Structured Protection of Stakeholder Interests
In this proof of concept, AILuminate completed reliability characterization of AI systems without exposing test data to developers or model weights to MLCommons or AVERI.
Moving Toward an Industry Standard
MLCommons has already applied similar methods to the most sensitive domain of health informatics, building Trusted Execution Environments (TEE) through the MedPerf project to enable joint evaluation of encrypted data across multiple parties without any information leakage.
In the future, AI adopters will be able to quickly run trusted evaluations tailored to specific deployments. Confidentiality guarantees for models, data, and tests—combined with fresh prompts supplied by independent institutions—will provide the clearest basis for decisions on AI adoption, deployment, and reliable operation.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接