MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs

MLCommons today announced the launch of MLPerf Mobile v6.0, adding generative AI benchmarks for running large language models (LLMs) on Android devices. These tests join existing benchmarks for image generation, object detection, and super resolution in the MLPerf Mobile app to form a complete test suite.

New On-Device LLM Benchmarks

MLPerf Mobile v6.0 adopts the following models as new LLM benchmarks:

  • Llama 3.2 1B Instruct
  • Llama 3.2 3B Instruct
  • Llama 3.1 8B Instruct

The models will process requests from the TinyMMLU and IFEval datasets to quantify the performance and accuracy of on-device AI inference.

LLM tests can run via CPU on devices with sufficient memory, without requiring custom accelerators. Additionally, this release supports NPU-accelerated execution of the Llama 3.1 8B Instruct model on devices equipped with the Qualcomm Snapdragon 8 Elite Gen 5 SoC. The working group plans to expand LLM acceleration support to more devices and platforms in the future.

Expanded SoC Support and Broad Availability

To quickly integrate support for new devices, v6.0 adds support for devices based on the MediaTek Dimensity 9500 series chips. It also updates support for the following chips:

  • Qualcomm Snapdragon 8 Elite Gen 5
  • Samsung Exynos 2600

The app now supports NPU-accelerated execution across a wide range of mobile devices.

The MLPerf Mobile app is available via the Google Play Store, Apple App Store, and the MLPerf Mobile GitHub repository. The GitHub repository also provides the complete open-source code under the Apache 2.0 license.

This article is from MLC blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!