LMSYS Org, in collaboration with MLCommons, announced the official launch of the Ares benchmark, the first open-source standardized framework in the AI industry dedicated to long-context multi-agent reasoning. This benchmark aims to address the shortcomings of existing evaluations in complex agent tasks, providing more reliable model performance metrics.
Core Design of Ares
Ares is built on the foundation of Chatbot Arena, incorporating an advanced Elo Rating system for dynamic model ranking. The test scenarios cover tool calls, multi-turn dialogues, and long-context understanding, totaling over 5,000 high-quality task datasets.
- Long-Context Reasoning: Supports up to 128K token input, simulating real-world agent applications.
- Multi-Agent Collaboration: Evaluates model coordination capabilities in team tasks.
- SGLang Integration: Leverages the SGLang framework for efficient inference, accelerating the benchmark by over 10x.
Initial Leaderboard Results
On the Ares leaderboard, top models have shown impressive performance:
- Claude 3.5 Sonnet: Elo 1452
- GPT-4o: Elo 1438
- Llama 3.1 405B: Elo 1395
- Gemini 1.5 Pro: Elo 1372
These scores are based on a combination of millions of user votes and automated evaluations, ensuring objectivity.
Open Source and Community Contribution
Ares is fully open-source, with code and datasets released on GitHub and Hugging Face. Developers can quickly get started via pip install ares-bench. MLCommons calls on the community to submit new tasks to drive benchmark iteration.
This release marks the evolution of AI evaluation from a single Chatbot Arena to a multi-agent ecosystem, helping to standardize the industry. (Full coverage of announcement highlights)
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接