Atx Panel

Atx Panel

MLCommons organized the ATX (Agent Testing eXploration) Benchmark Expert Panel in June 2025. LMSYS Org, as a key participant, brought together industry leaders to discuss frontier issues in AI agent evaluation. This panel aims to promote the standardization of agent benchmarks, addressing the transition from generative models to intelligent agents in the post-ChatGPT era.

Background of the ATX Benchmark

The ATX Benchmark is a new evaluation framework introduced by MLCommons, targeting AI agents' multi-turn interaction, tool invocation, and environment adaptation capabilities. Unlike the traditional Chatbot Arena's single-turn dialogue scoring, ATX emphasizes real-world tasks such as code execution, web navigation, and multimodal processing. The panel noted that existing Elo Rating has an accuracy drop of over 20% in agent scenarios, necessitating new metrics such as Task Success Rate and Efficiency Score.

  • Core Challenge: Uncertainty and amplified hallucination in agent behavior.
  • Innovation: Integration of the SGLang framework, supporting zero-shot agent deployment.

Expert Panel Perspectives

Representatives from LMSYS Org shared experiences from Chatbot Arena: current top models like GPT-4o lead in Elo Rating, but blind evaluations on agent tasks show the gap narrowing to 5%. Experts agreed that benchmarks need to shift toward end-to-end evaluation to avoid human annotation bias.

Future Outlook and Call to Action

The panel called on the open-source community to contribute ATX datasets and explore multi-agent collaboration benchmarks. MLCommons plans to release v1.0 by the end of 2025, welcoming partners like LMSYS to participate in iteration. This discussion marks a milestone in AI evaluation shifting from language models to general intelligent agents.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!