A new GPT-OSS benchmark and DeepSeek R1 updates for latency-optimized reasoning

The MLPerf Inference v6.0 release marks a significant expansion in open-source large language model (LLM) coverage, introducing two key additions to the Reasoning LLM task group: the GPT-OSS 120B benchmark based on a high-capacity MoE model, and a new interactive workload for DeepSeek-R1 with low-latency constraints, featuring the first standardized speculative decoding in MLPerf.

MLC MLPerf Inference GPT-OSS 120B
1,012