GPT-OSS 20B: A Sparse MoE Pretraining Benchmark for MLPerf Training v6.0
GPT-OSS 20B is a new Mixture-of-Experts (MoE) pretraining benchmark introduced in MLPerf Training v6.0, designed to lower the barrier to entry and evaluate sparse architectures. It features variance reduction techniques to ensure fair comparisons and can run on a single 8-GPU node.