MLPerf Training Introduces Its First LLM Post-Training Benchmark

MLPerf Training Introduces Its First LLM Post-Training Benchmark

Starting with its v6.1 submission round in October 2026, MLPerf Training will add an LLM post-training benchmark, complementing the existing pre-training benchmark suite. Pre-training builds foundational intelligence from vast amounts of data, while post-training refines models so they perform better on specific tasks. Since the second half of 2025, major advances in LLM training have come largely from post-training scaling — for example, small models approaching frontier performance, or generational improvements achieved on the same base model.

Figure 1

The MLPerf Reasoning Task Force designed a benchmark capable of completing a real LLM post-training workload within 1024 GB300 hours. Real post-training combines distillation, reinforcement learning, and supervised fine-tuning, spanning math, reasoning, and agentic tasks. The task force ultimately chose reinforcement learning with verifiable rewards (RLVR), which suits software task scenarios. An LLM-driven coding agent solves agentic software engineering (SWE) tasks in a sandboxed environment, and each attempt is called a rollout. The system samples multiple rollouts per problem and assigns a binary pass/fail reward, then updates the model using Group Relative Policy Optimization (GRPO).

Model Selection

The benchmark uses the Qwen 3.5 397B model released in February 2026, the largest open-weight model in the Qwen 3.5 series by parameter count, under the Apache 2.0 license. The model has a mixture-of-experts (MoE) architecture with 397 billion total parameters and 17 billion activated parameters per token. It is the first to productize a hybrid architecture combining Gated DeltaNet (GDN) with sparse MoE, activating 1 shared expert and 10 of 512 routed experts (2%) in each MoE layer. GDN replaces the linearly growing KV cache with a fixed-size compressed state, making it particularly well suited to long-context scenarios.

Dataset Selection

The benchmark uses the R2E-Gym SWE problem dataset under the Apache 2.0 license, containing 700 training problems and 251 validation problems generated from GitHub commits across multiple Python projects. The agent receives the problem description in a Linux container with dependencies pre-installed and must produce a fix patch. The task force filtered the dataset for difficulty and fixed cross-architecture compatibility issues.

Environment and Implementation

The benchmark caps the maximum context at 65536, a maximum of 30 agent turns, and 16 generations per prompt. The reference implementation is based on NVIDIA NeMo-RL, using Ray to coordinate policy training and rollout generation, supported by Megatron Core, vLLM, and the OpenHands harness respectively. Evaluation uses test files that are invisible to the agent to prevent reward hacking.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!