MLCommons today announced the latest results of the MLPerf Training v6.0 benchmark suite. The two new benchmarks added in this round and the large number of submissions highlight the rapid transformation of the AI ecosystem.
"This is an exciting time for the community," said Shriya Rishab, co-chair of the MLPerf Training working group. "We see convergence in best practices for AI model training, while the technical diversity of underlying frameworks and systems is also increasing."
New Benchmarks Emphasize Sparse Computing
The MLPerf Training benchmarks cover models, software, and hardware through full system tests. v6.0 adds two new benchmarks, DeepSeek V3 and GPT-OSS 20B, both using Mixture-of-Experts (MoE) architecture, reflecting the industry's shift toward sparse computing.

DeepSeek V3 has 671 billion total parameters, activating 37 billion parameters per token, making it the largest benchmark in the suite. GPT-OSS 20B, with 21 billion total parameters and 3.6 billion activations per token, is smaller and suitable for single 8-GPU node testing.
Record Diversity in Submitted Systems
v6.0 received 95 unique systems, involving 13 hardware accelerators, 19 host processors, with 60% being multi-node systems. The number of cloud systems more than doubled compared to v5.1.

Submitters used various FP4 precision schemes, highlighting the industry's exploration of low-precision training. MLPerf's accuracy threshold requirements help the industry clearly compare the performance differences between different implementations.
24 Organizations Participate, Ecosystem Continues to Grow
The results come from 24 organizations including AMD, NVIDIA, Google, Azure, among which Inventec, Netweb Technologies India LTD, TTA, and Vultr are first-time submitters. MLCommons welcomes more organizations to join the working group to jointly improve the benchmarks.
The full results are available on the MLCommons official website.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接