MLCommons released the latest results of the MLPerf Inference v6.1 benchmark. This release set a new high in the number of submitting organizations, introduced two new tests aligned with recent AI inference deployment trends, and for the first time published peer-reviewed performance results for several new AI platforms, with up to 5.7x higher peak performance than a year ago.
Benchmark Evolution: New Tests and Optimized Support
MLPerf Inference v6.1 introduces two new tests that reflect the industry's evolution toward more complex, multi-step, and agentic AI inference deployments, covering data center and edge scenarios.
End-to-End Retrieval-Augmented Generation (RAG) Benchmark
This test evaluates question-answering performance of complex multi-model pipelines, including steps such as embedding models, retrievers, rerankers, and LLMs, separately measuring document corpus ingestion and vector-database-based query answering tasks.
Edge Agentic Inference Benchmark
This test focuses on the shift from single-turn interaction to complex multi-turn agentic workloads, such as agentic coding. It uses edge models and quantization, supports single-stream coding workloads and latency metrics, and includes accuracy measurement under time constraints as well as deterministic workload performance testing.
In addition, the benchmark adds support for Speculative Decoding optimization, a technique that can predict and verify multiple tokens in a single forward pass and has been applied to interactive scenarios and GPT-OSS tasks.
Submission Results Highlight the Pace of AI Innovation
This benchmark received submissions from 30 organizations, a new record, including AMD, NVIDIA, Intel, and others. New hardware includes AMD Ryzen AI Max+ 395, AMD Instinct MI350P, Intel Arc Pro B70, and the previewed NVIDIA Rubin and Vera Rubin NVL72. The largest system contains 512 accelerators, and the first heterogeneous system also appeared.
Performance continues to break records: visual language model (VLM) server scenario per-accelerator results improved 2.99x over v6.0; Deepseek R1 test improved 5.7x over v5.1.
Broad Industry Participation and Future Outlook
First-time submitters include six organizations such as Atlas Inference and Crusoe. More than 50% of submitters adopted MLPerf's new API-centric harness, laying the foundation for the next-generation MLPerf Endpoints benchmark.
Detailed results are available on the MLCommons website and visualization dashboard.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接