Whether AI systems are applied in healthcare, banking, or energy sectors, their complex black-box nature requires us to first clarify expected behaviors and then evaluate the system's delivery reliability. Only by understanding reliability can we manage risks (e.g., whether 80% reliability is sufficient for life-critical systems) and control costs (e.g., the acceptable level of customer service errors under cost reduction).
Measuring AI software reliability is the core mission of the MLCommons AI Risk and Reliability Working Group. Improving AI reliability is crucial for market growth and societal protection. We believe that successfully enhancing AI reliability within the industry requires a systematic and thoughtful plan, along with effective implementation, iteration, and maintenance of that plan.
Like any long-term roadmap, our starting point is to create a map—this AI Reliability Map will evolve and improve over time.
We first focus on pre-deployment testing of AI system behavioral reliability. Within the AI application lifecycle, reliability must be addressed during development, deployment, and operations: process in development, testing in deployment, and monitoring in operations. We initially focus on deployment testing because it offers the most concrete opportunities for change, particularly regarding AI system behavior: its responses and actions taken. Established methods already exist for hardware and conventional software reliability management.
At the core of AI Reliability (AIR) is consistently following behavioral rules across different environments. We introduce the AI Reliability Map to connect these issues with the fundamental concept of consistently following rules across different environments:
AI Reliability Map: Rules and EnvironmentsEnvironmentCorrectness: Following rules under instruction complianceSafety: Preventing rule violations caused by malicious actorsRulesFunctionalityNeeds testingNeeds testingData ProtectionNeeds testingNeeds testingProduct SafetyNeeds testingNeeds testingFrontier SafetyNeeds testingNeeds testingPsychosocial LimitsNeeds testingNeeds testingThe rows in the matrix above represent the rules the system must follow, and the columns represent the environments in which the system must follow those rules. Whether under normal instructions or malicious behavior (such as prompt injection or misinformation), functional rules must be followed. Similarly, normal instructions may entice the system to violate privacy rules, posing risks as dangerous as attacks. This is why we must follow all rules in all environments, from normal use to malicious behavior.
Notably, AI safety cannot be tested in isolation: testing must attempt to violate rules, and system behavior may vary significantly depending on which rule is tested, so testing must be conducted across all system behaviors.
This AI Reliability Map helps us define the problem, but taking action requires further detail. Below are more detailed categories and subcategories for rules and environments. These subcategories are designed to address prominent known issues in pre-deployment testing of commercial systems. The entire matrix is extensible to accommodate emerging or increasingly important issues.
Notable points about this expansion: First, functionality covers compliance with regulations and deployment requirements, with the former always taking precedence over the latter. Second, data protection covers personal data privacy expectations and "discrete information management," such as appropriate use of corporate data and intellectual property. Third, in our map, frontier safety covers CBRN and offensive cyber.
The color blocks indicate the current state of AI testing. Yellow represents the scope of most public capability tests. Green and blue represent the scope of the MLCommons AILuminate safety and jailbreak benchmarks.
Any mid-to-high-level AI agent with a natural interface, regardless of its purpose, could theoretically fail across the entire range. A personal finance AI system might provide incorrect financial advice, design a virus on request, enable hackers to access confidential company information, or trick users into unintended purchases. These AI tools face both domain-specific risks (e.g., incorrect financial advice, unintended purchases) and general risks (e.g., viruses, hacking) across various verticals. Our challenge, as an industry and field, is to develop a well-structured yet evolving deployment testing methodology that covers this map.
This is the work of the MLCommons AI Risk and Reliability Working Group, where industry, academia, and government collaborate to translate such frameworks into actionable benchmarks, including AILuminate. As an open engineering consortium, MLCommons is uniquely positioned to lead this effort—uniting organizations that build AI systems with those that deploy, regulate, or are affected by them. If your organization is working to understand or improve AI reliability, we hope to build together with you. Learn more and join the AIRR Working Group at mlcommons.org/working-groups/ai-risk-reliability.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接