MLCommons Joins EU-Funded AIRIS Project to Build and Benchmark Next-Generation Biomedical AI

MLCommons Joins EU-Funded AIRIS Project to Build and Benchmark Next-Generation Biomedical AI

This summer, MLCommons officially became a member of the AIRIS project consortium (Mechanism-Informed Multimodal Generative AI for Causal and Dynamical Modeling in Biomedical Research), funded by the EU Horizon Europe program. The project comprises 21 institutions from Europe, Canada, and the United States, and will receive €16.9 million in funding over the next four years.

Project Goals and Core Directions

AIRIS aims to develop generative AI models that integrate biological mechanism knowledge with clinical data, creating an "AI collaborator" that can assist researchers in understanding disease progression and advancing personalized medicine. The platform will construct and reason over disease mechanism models rather than relying solely on statistical patterns, helping to discover unknown disease pathways and propose new scientific hypotheses.

MLCommons' Core Role

As a global leader in AI benchmarking, MLCommons will be responsible for developing comprehensive evaluation frameworks and integrated benchmark suites, conducting rigorous independent assessments of the platform across five dimensions: accuracy, robustness, fairness, interpretability, and usability. The benchmarks will cover five disease areas and specifically detect bias across patient subgroups including gender, ethnicity, and age.

Alexandros Karargyris, head of the MLCommons medical working group, said: "Rigorous independent evaluation can transform promising AI systems into tools that researchers truly trust. We will build open benchmarks that examine not only accuracy but also fairness, robustness, and interpretability in real disease scenarios."

Future Plans

MLCommons will coordinate AIRIS's iterative evaluation rounds, track progress across platform versions, and open the benchmarks to external researchers, enabling the scientific community to test AIRIS on their own data and ultimately establishing it as a reference benchmark for multimodal, mechanism-driven generative AI in biomedical research.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!