Technical Standards are the Bridge to Enabling AI Adoption

As artificial intelligence transitions from entertaining consumer chat experiences to a general-purpose technology driving enterprise services in the economy, it faces significant reliability hurdles. Enterprises need to trust that AI systems can produce correct, safe, and reliable responses before placing them in roles where they can create greater value.

To build the trust required for widespread enterprise adoption, the industry must adopt risk management standards that reduce uncertainty for deployers. Until enterprises—including small and medium-sized businesses—feel comfortable allowing AI agents to access their corporate data and autonomously negotiate pricing agreements, we cannot achieve the automated transactions that the industry now pursues. What level of reliability must an AI vision system demonstrate—how many 9s after 99.9…%—to be trusted for inspecting oil pipeline damage? What requirements must be met for deploying AI clinical support tools to assist doctors in diagnosis? In manufacturing lines, where one hour of downtime can cost millions in revenue, what should deployment standards be? Deploying AI systems in high-risk, high-trust applications such as finance, healthcare, and manufacturing will require far higher levels of reliability than today. This also means we need to reliably measure that reliability.

Ultimately, reliability targets and procedural requirements are set by consensus standards such as ISO/IEC 42001, just like other industries that require risk management. Since AI is a probabilistic technology, evaluation standards that support these targets are essential for continuously and empirically demonstrating reliability and compliance.

AI's probabilistic nature makes it fundamentally different from other technologies. For example, a civil engineer can sign off on a bridge design that meets standards, almost fully confident that it can carry vehicles in various weather conditions, because the bridge does not change after the hundredth vehicle passes. An LLM, however, produces different results with each interaction. This probabilistic behavior gives the new technology powerful adaptability, but also makes reliable measurement and evaluation extremely difficult.

Therefore, AI developers need to apply the same scrutiny when designing systems—reviewing plans and ensuring they meet objectives—and at the same time, must continuously measure and empirically demonstrate compliance with reliability targets under diverse real-world conditions. By design, AI generates different outputs even when given the same input twice, so it is necessary to empirically measure model inputs and outputs in different contexts to determine whether risks are adequately mitigated.

MLCommons' Role

This is where we come in. Technical standards organizations like MLCommons are a vital complement to traditional standards bodies such as ISO. Standards set by organizations like ISO establish broad direction, clear objectives, and qualitative requirements based on business needs and societal concerns. Benchmark standards organizations then translate these objectives into precise, actionable metrics. This relationship ensures that the objectives in ISO standards are grounded in empirical data that model developers and enterprise users can practically apply.

For example, MLCommons actively participates in ISO work such as the 42119 series (AI testing and assurance standards). The industry needs broad, international consensus-driven guidelines for AI measurement, which are then implemented through concrete benchmarks like MLCommons AILuminate for generative AI safety and product reliability. These technical specifications must evolve rapidly to match the pace of AI innovation, providing a "living" bridge between standard objectives and industry practice.

Standardized Evaluation Drives Progress

Ultimately, standardized evaluation drives progress and builds public trust. Historical precedents like the New Car Assessment Program (NCAP) show that rigorous safety testing can transform an entire industry, raising the market share of five-star safety-rated vehicles from negligible to over 86% in large markets. By applying similar technical rigor to AI, and through evolving benchmarks like AILuminate, the industry can ensure AI is safer and more reliable, unlocking higher-value markets for companies and greater value for consumers.

Join the Effort

Building trustworthy AI requires global collaboration. Join MLCommons to help shape the technical standards that will define AI reliability for the next decade. With over 125 member organizations already contributing to benchmarks like AILuminate, every organization committed to making AI safer, more reliable, and broadly trusted has a seat at the table.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!