Ailuminate Jailbreak V05

Ailuminate Jailbreak V05

Introduction: AILuminate Jailbreak V05 Fully Upgraded

MLCommons, in collaboration with LMSYS Org, has launched the AILuminate Jailbreak V05 benchmark, the latest standard for evaluating large language model (LLM) resistance to jailbreak attacks. This version focuses on high-risk scenarios, including chemical weapon synthesis, biological toxin manufacturing, and cyber intrusion, introducing more complex multi-turn attack chains and roleplay prompts. Through thousands of human evaluations, each model's jailbreak resistance Elo rating is calculated, similar to the scoring mechanism of Chatbot Arena.

Testing Methodology and Innovations

  • Attack Dataset: Expanded to 200+ jailbreak prompts covering 8 major danger categories, optimized using automated generation tools.
  • Inference Framework: Integrates SGLang for efficient multi-turn inference, supporting long-context attacks.
  • Evaluation Protocol: Human judges anonymously compare model output safety, with win rates converted to Elo scores. Confidence intervals based on at least 64 matchups.
  • New Features: Introduces 'roleplay jailbreak' and 'code injection' variants to simulate real-world attack paths.

Leaderboard Highlights: Claude Leads, GPT Close Behind

On the V05 leaderboard, Claude 3.5 Sonnet tops with 1485 Elo, demonstrating outstanding safety alignment. Anthropic's Claude 3 Opus (1462) and OpenAI's GPT-4o (1472) rank second and third. Among open-source models, Meta's Llama 3.1 405B scores 1421, significantly ahead of Mistral Large 2's 1378.

  • Top 5:
    1. Claude 3.5 Sonnet: 1485 ± 12
    2. GPT-4o: 1472 ± 11
    3. Claude 3 Opus: 1462 ± 13
    4. Llama 3.1 405B: 1421 ± 15
    5. GPT-4o-mini: 1405 ± 14

Lower-tier models such as Gemini 1.5 Pro score only 1038, exposing the vulnerability of lightweight LLMs.

Key Insights and Model Comparison

V05 results show that jailbreak resistance is highly correlated with general capabilities (correlation coefficient 0.92), but not absolute: some instruction-tuned models lag in safety. Claude series benefits from constitutional AI training, while GPT-4o excels in multi-turn defense. Open-source models have made significant progress but still require strengthened post-training safety mechanisms.

ModelElo RatingChange (vs V04)
Claude 3.5 Sonnet1485+23
GPT-4o1472+15
Llama 3.1 405B1421+45

Conclusion and Outlook

AILuminate V05 highlights the intensity of the AI safety race, calling on developers to prioritize investment in defense mechanisms. Future versions will incorporate more real-world attacks and explore multimodal jailbreaks. Visit the full leaderboard: MLCommons Official Website.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!