Ares Announce

Ares Announce

LMSYS Org, in collaboration with MLCommons, announced the official launch of the Ares benchmark, the first open-source standardized framework in the AI industry dedicated to long-context multi-agent reasoning. This benchmark aims to address the shortcomings of existing evaluations in complex agent tasks, providing more reliable model performance metrics.

Core Design of Ares

Ares is built on the foundation of Chatbot Arena, incorporating an advanced Elo Rating system for dynamic model ranking. The test scenarios cover tool calls, multi-turn dialogues, and long-context understanding, totaling over 5,000 high-quality task datasets.

  • Long-Context Reasoning: Supports up to 128K token input, simulating real-world agent applications.
  • Multi-Agent Collaboration: Evaluates model coordination capabilities in team tasks.
  • SGLang Integration: Leverages the SGLang framework for efficient inference, accelerating the benchmark by over 10x.

Initial Leaderboard Results

On the Ares leaderboard, top models have shown impressive performance:

  • Claude 3.5 Sonnet: Elo 1452
  • GPT-4o: Elo 1438
  • Llama 3.1 405B: Elo 1395
  • Gemini 1.5 Pro: Elo 1372

These scores are based on a combination of millions of user votes and automated evaluations, ensuring objectivity.

Open Source and Community Contribution

Ares is fully open-source, with code and datasets released on GitHub and Hugging Face. Developers can quickly get started via pip install ares-bench. MLCommons calls on the community to submit new tasks to drive benchmark iteration.

This release marks the evolution of AI evaluation from a single Chatbot Arena to a multi-agent ecosystem, helping to standardize the industry. (Full coverage of announcement highlights)

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!