Since its launch in 2018, MLPerf has been the industry standard for measuring AI system performance. Over this period, MLPerf has documented more than 100x improvement in performance per watt for large language model inference and more than 50x improvement in training speed. Over the past eight years, the AI industry has matured, and AI services are now used daily by enterprises and consumers worldwide.
The Background Behind MLPerf Endpoints
MLPerf was originally designed to help a small number of cloud providers and system builders make AI hardware procurement decisions. Today, procuring inference compute has become a critical business decision for enterprises of all sizes, requiring simultaneous evaluation of new clouds, cloud service providers, and managed services. These buyers need reliable, comparable, and independent performance benchmarks, and the benchmarks must dynamically keep pace with an industry that releases new models on a weekly basis.
Four Core Principles
MLPerf Endpoints is designed to meet this diverse and complex need, developed around four buyer-centric principles:
- Current: Results stay in sync with the market—buyers don't need to wait months for new hardware or models to be included.
- Comprehensive: Covers numerous competitive inference providers, systems, and workloads.
- Comparable: Provides apples-to-apples comparisons across vendors, with support for cost or power normalization.
- Commentary: Delivers additional decision context through visualization, data filtering, and analysis.
v0.7 Release Highlights
MLPerf Endpoints v0.7, as the foundational release, has published initial results from Coreweave, Google, Intel, KRAI, and Nvidia, covering three benchmarks with performance spanning multiple orders of magnitude. The release supports an automated submission pipeline, continuous review tooling, and dynamic result visualization, viewable at mlcommons.endpoints.com. The rules are evolving toward a more buyer-centric direction.
Future Plans and Participation
A v1.0 release will arrive later this year, featuring more buyer-centric rules, normalization processing, and expanded benchmarks (including agentic workloads), while opening a rolling submission process to a broader set of members. MLCommons thanks its 30+ supporters, including AMD, Argonne National Laboratory, Broadcom, and others. System or service providers can participate in rule-making and submit immediately; enterprise buyers can provide feedback on requirements via email.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接