The key to trustworthy AI evaluation is secrecy by design

The key to trustworthy AI evaluation is secrecy by design

Every enterprise deploying AI wants confirmation that the system runs reliably and safely under its specific use cases and data. But simply connecting and running tests is far from straightforward.

Privacy and Integrity: A Dual Challenge

Deployers such as banks must protect sensitive data, while frontier labs strictly guard model weights. More critically, benchmark providers must have both technical and legal measures in place to prevent models from contaminating benchmarks through learning. Confidentiality of the evaluation alone is not sufficient; robust benchmark management procedures are also essential.

First Proof of Concept for Double-Blind Evaluation

To demonstrate top-tier benchmark integrity, MLCommons partnered with Google DeepMind, OpenMined, and AVERI to conduct the first double-blind evaluation of closed-source models. This approach leverages cryptographic safeguards to ensure evaluation components remain fully usable without contaminating the benchmark.

A held-out subset of the AILuminate™ safety benchmark was used for this validation, ensuring the Google DeepMind model had never previously encountered the evaluation set. AVERI leveraged OpenMined's secure computation to run prompts in a containerized environment, providing provable protection for both model weights and the evaluation itself.

Structured Protection of Stakeholder Interests

In this proof of concept, AILuminate completed reliability characterization of AI systems without exposing test data to developers or model weights to MLCommons or AVERI.

Moving Toward an Industry Standard

MLCommons has already applied similar methods to the most sensitive domain of health informatics, building Trusted Execution Environments (TEE) through the MedPerf project to enable joint evaluation of encrypted data across multiple parties without any information leakage.

In the future, AI adopters will be able to quickly run trusted evaluations tailored to specific deployments. Confidentiality guarantees for models, data, and tests—combined with fresh prompts supplied by independent institutions—will provide the clearest basis for decisions on AI adoption, deployment, and reliable operation.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!