MLCommons Agent Reliability Profile Named a Finalist in Global Agentic Regulator Hackathon

MLCommons Agent Reliability Profile Named a Finalist in Global Agentic Regulator Hackathon

How can we determine whether an AI agent will strictly adhere to its permitted scope? This is a question financial institutions and their regulators currently struggle to answer with confidence, and it is the core topic of the Agent Reliability Profile introduced by the MLCommons Financial Services Working Group. The framework has successfully been selected as a finalist in C:>DIR’s global “Agentic Regulator” Hackathon.

Why It Matters

AI agents can not only answer questions but also perform actual operations. This makes them powerful, but it also requires institutions and regulators to find reliable ways to specify the scope of operations agents are permitted to perform and whether they truly comply with those limits. Current agent identity standards can prove “who” an agent is, but cannot specify what it “should be authorized to do” once connected to financial digital infrastructure. This gap is called the authorization and oversight gap.

The Agent Reliability Profile is the working group's flagship project. It aims to fill this gap through a standardized, regulator-compatible framework for describing, validating, and benchmarking the reliability of agent deployments in financial services.

C:>DIR Hackathon

C:>DIR is hosted by the University of Cambridge's Digital Innovation and Regulation Initiative and supported by more than 35 organizations, including the BIS Innovation Hub, the Global Financial Innovation Network (GFIN), and the Digital Regulation Cooperation Forum. It is a global competition for regulators and industry aimed at building practical, deployable prototypes to enhance trust and accountability mechanisms for agentic AI.

The scale of the competition shows the level of industry attention to this issue: it received 336 submissions from more than 65 countries, and 36 teams (six per problem area) ultimately advanced to the finals.

Our Submission

The submission by working group co-chairs Mike Hsu and Medha Bankhwal competed in the Know Your Agent (KY-A), Digital Verification, and Digital Public Infrastructure track. The submission demonstrates two tools built on the Agent Reliability Profile:

  • Profile Builder: reads an institution's own evidence (design documents, configuration exports, policies), compiles it into a Level 1 assertion profile, and highlights discrepancies between stated intent and the actual built configuration.
  • Profile Validator: uses the profile as a test specification, runs falsification tests against real system behavior, and ultimately produces a Level 2 validation profile.

The demonstration applies these two tools in an open banking data sharing and consent scenario, testing whether agents strictly adhere to the scope of data sharing that customers actually consented to. This is a real, high-stakes test of the authorization and oversight gap.

Next Steps

The finalist build phase runs from September 1 to 8. Live demonstrations and judging by a panel of regulatory researchers from the Cambridge Regulatory Visiting Scholars Programme will take place on September 15-16. The winners will be announced on September 18, 2026, at the C:>DIR Summit.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!