MLCommons Releases New MLPerf Inference v6.0 Benchmark Results

MLCommons® recently announced the release of the latest results of its industry-standard MLPerf® Inference v6.0 benchmark suite. This update introduces several major advancements, ensuring the benchmark covers real-world scenarios of current AI deployments and comprehensively demonstrates AI system performance.

In the 11 data center tests of MLPerf Inference v6.0, five are new or updated, while edge systems have added object detection tests. Key changes include:

  • A new open-source large language model benchmark based on GPT-OSS 120B, supporting math, scientific reasoning, and coding tasks;
  • An expanded advanced reasoning benchmark for DeepSeek-R1, adding interactive scenarios with speculative decoding support;
  • DLRMv3, the third generation of the recommendation system benchmark, introducing sequential recommendation testing for the first time, with significant engineering contributions from Meta;
  • The suite’s first text-to-video generation benchmark;
  • A new vision-language model (VLM) benchmark that transforms multimodal data from the Shopify product catalog into structured metadata;
  • An upgraded edge single-shot object detection benchmark based on the Ultralytics YOLOv11 Large model.

“This is our most significant revision to the Inference benchmark suite,” said Frank Han, Systems Development Engineering Technologist at Dell Technologies and co-chair of the MLPerf Inference Working Group. “The enthusiastic collaboration and engineering contributions from members are unprecedented, driving us to update multiple benchmarks to keep pace with the rapid evolution of AI models and technologies, ensuring test relevance and representativeness.”

The open-source MLPerf Inference benchmark suite measures system performance in an architecture-neutral, representative, and reproducible manner, aiming to provide a fair platform for industry competition and promote innovation, performance, and energy efficiency improvements. The published results offer critical technical information for customers procuring and tuning AI systems.

“We thank Meta, Shopify, and Ultralytics for their deep collaboration in providing datasets, task definitions, and expertise,” said Miro Hodak, Senior Technical Staff at AMD and co-chair of the MLPerf Inference Working Group. “These partnerships ensure the tests reflect the latest state of the industry.”

“MLPerf Inference benchmarks drive transparency and accountability in the AI industry,” said Glenn Jocher, CEO and founder of Ultralytics. “We use them to validate the real-world performance of YOLO models, helping developers make informed decisions.”

New Tools for Submitters and Users

Inference 6.0 introduces a new harness, LoadGen++, allowing LLMs to run with the typical serving-style software stack used today. “LoadGen++ is a major upgrade over the previous generation, helping us agilely track state-of-the-art technology,” Han added.

Additionally, results are now viewable on a new online dashboard on the MLCommons website, supporting advanced filtering and custom performance charts: https://mlcommons.org/visualizer.

Large-Scale Multi-Node Systems in Spotlight

Inference 6.0 submissions show that technology providers are eager to demonstrate the performance of multi-node systems on real inference workloads. Multi-node submissions increased by 30% compared to Inference 5.1 six months ago, with 10% of systems exceeding 10 nodes (only 2% in the previous round), and the largest system reaching 72 nodes with 288 accelerators, four times the node count of the largest system in the previous round.

“As AI applications enter production, demand for large-scale high-performance systems has surged,” said Hodak. “Multi-node systems introduce unique challenges including architecture, networking, storage, and software optimization, and stakeholders are actively addressing large-scale inference.”

AI Community Continues to Embrace MLPerf Inference

This benchmark round received submissions from 24 organizations: AMD, ASUSTeK, Cisco, CoreWeave, Dell, GATEOverflow, GigaComputing, Google, Hewlett Packard Enterprise, Intel, Inventec Corporation, KRAI, Lambda, Lenovo, MangoBoost, MiTAC, Nebius, Netweb Technologies India Limited, NVIDIA, Oracle, Quanta Cloud Technology, Red Hat, Stevens Institute of Technology, and Supermicro.

“Welcome to first-time submitters Inventec Corporation, Netweb Technologies India Limited, and Stevens Institute of Technology,” said Han. “Thank you to members, contributors, and partners like Meta, Shopify, and Ultralytics for building together the most comprehensive AI inference performance benchmark, helping the community make better decisions.”

View Results

Visit the MLPerf Inference v6.0 Results Dashboard for details.

About MLCommons

MLCommons is the global AI benchmark leader, an open-source engineering alliance supported by over 130 members, bringing together academia, industry, and civil society to drive AI measurement and improvement. Since the launch of the MLPerf benchmark in 2018, it has become the industry standard for machine learning performance, promoting transparency, security, speed, and efficiency. For more information, visit MLCommons.org or email inquiries.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!