MLCommons Lays the Foundation for Defensible Jailbreak Benchmarking
MLCommons introduces a taxonomy-based methodology for benchmarking single-turn jailbreak attacks on large language models, establishing a defensible and reproducible evaluation framework that prioritizes structural coverage over scale.