Aaai2025

Aaai2025

LMSYS Org, a key player in the AI community, unveiled the latest benchmark results of Chatbot Arena at the AAAI 2025 conference. This update not only refreshes the global AI chatbot leaderboard but also provides developers with valuable insights for model optimization.

Chatbot Arena Benchmark Overview

Chatbot Arena is a pioneering platform introduced by LMSYS that generates Elo Rating scores through anonymous user-vs-model voting. The system simulates real-world scenarios, allowing users to blindly evaluate responses from different models, ultimately producing an authoritative ranking. As of this update, it has accumulated over 3 million votes covering more than 100 models.

Top Model Performance

  • Claude 3.5 Sonnet: Elo 1308, dominating the leaderboard for consecutive months, particularly excelling in complex reasoning and creative tasks.
  • GPT-4o: Elo 1302, highly balanced with leading multimodal capabilities.
  • Gemini 1.5 Pro: Elo 1290, outstanding long-context processing abilities.
  • Open-source highlight: Llama 3.1 405B Elo 1285, cost-effective, narrowing the gap with closed-source models.

Technological Innovations such as SGLang

The report specifically highlights SGLang, an efficient inference framework that can increase model throughput by 2-5x. Through RadixAttention and zero-overhead batching, SGLang significantly reduces latency and supports real-time multi-turn conversations. The LMSYS team demonstrated its application in Arena, helping models maintain high Elo scores under heavy load.

Key Data Comparison

ModelElo RatingWin Rate (%)Category Advantage
Claude 3.5 Sonnet130858.2Reasoning/Coding
GPT-4o130257.5Multimodal
Llama 3.1 405B128555.1Open-source/Cost

Industry Impact and Outlook

This AAAI 2025 update highlights the importance of user-driven evaluation, avoiding biases common in traditional benchmarks. LMSYS Org calls for more models to join Arena to foster the development of the open-source ecosystem. In the future, they plan to integrate more Chinese and multilingual tests to promote fair global AI competition.

For more details, visit the original link.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!