Where the Industry Is Investing: A Look at MLPerf Inference v6.1

Where the Industry Is Investing: A Look at MLPerf Inference v6.1

Each round of MLPerf Inference results is a real-time snapshot of industry engineering investment. v6.1 stands out in two ways: a record number of submitters, and a clear shift toward agentic and end-to-end benchmarks, accompanied by a surge of new hardware.

Figure 1

Submitters and Submission Scale

This round has 30 submitters, spanning silicon vendors, OEM/ODM system vendors, cloud and neocloud providers, and specialized inference software companies. Some results are joint submissions, including Dell_AMD, Dell_MangoBoost, RedHat_Intel, and RedHat_Supermicro. A total of 120 systems were submitted, covering the Datacenter and Edge suites and both Closed and Open divisions.

Figure 2

Benchmark Overview

MLPerf Inference v6.1 includes 10 Datacenter benchmarks and 6 Edge benchmarks, with End-to-End RAG and Agentic Edge Inference as new additions. VLM adds an Interactive scenario, and the Interactive scenario for GPT-OSS-120B supports speculative decoding. The most popular Datacenter benchmark has become gpt-oss-120b, replacing Llama2-70b, which had led for several consecutive rounds, indicating significantly increased community acceptance of MoE models.

Figure 3

Performance Improvement Analysis

Compared with v6.0, VLM and DeepSeek R1 achieved the largest performance gains in the Offline and Server scenarios, mainly from NVIDIA Vera Rubin preview systems. Other benchmarks achieved incremental improvements on the same hardware through software and algorithmic optimizations.

Figure 4

Since its introduction in v4.0, Llama2-70b has run for 6 rounds, with cumulative median per-accelerator performance in the Server scenario improving 5.58x. Key drivers include FP4 low-precision computing, next-generation accelerators, and continuous software stack optimization.

Figure 5

The DeepSeek R1 benchmark improved Offline performance by 2.7x in one year, Server scenario by 5.7x, and the Interactive scenario also achieved 2.7x growth.

Notable Results

The number of multi-node submissions continued to rise, reaching a record high of 16 this round. Meanwhile, the scale of the largest submitted system increased from 288 accelerators in the previous round to 512 accelerators.

Figure 6Figure 7
This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!