Whisper Inferencev5 1
MLCommons has released the Whisper Inference v5.1 benchmark, introducing the large-v3 model and providing standardized performance data for speech-to-text model inference, aiming to help developers optimize deployment.
MLCommons has released the Whisper Inference v5.1 benchmark, introducing the large-v3 model and providing standardized performance data for speech-to-text model inference, aiming to help developers optimize deployment.
MLCommons released the Small LLM Inference 5.1 benchmark, evaluating SLM performance in real-world scenarios such as chatbots and text generation. Key updates include the addition of Llama 3.2 1B, optimized test scenarios, and expanded hardware support.
The latest report from LMSYS Org shows that DeepSeek Inference 5.1 stands out in the MLCommons September 2025 inference benchmark. Designed for large language models, it focuses on low-latency and high-throughput inference optimization.
MLCommons announces the release of MLPerf Inference v5.1 benchmark results, featuring new generative AI workloads and expanded model coverage, with NVIDIA, AMD, and others competing on latency and throughput.
The MLCommons organization has released the MLPerf Tiny v1.3 benchmark results, a key milestone for edge AI. It evaluates model performance on resource-constrained microcontrollers and embedded devices.
MLPerf Tiny v1.3 is the latest edge AI benchmark version from MLCommons, designed for resource-constrained devices. It introduces two new benchmarks (Image Classification and Visual Wake Words) and optimizes existing ones, with evaluation rules covering closed and open divisions.
MLCommons has officially launched Croissant MCP, a major upgrade to the Croissant metadata format designed specifically for AI model cards. This standard, contributed by partners including LMSYS Org, aims to address the fragmentation of current model documentation.
MLCommons and LMSYS Org have launched the AILuminate Jailbreak V05 benchmark, introducing complex multi-turn attack chains and roleplay prompts. Claude 3.5 Sonnet tops the ranking with 1485 Elo, followed closely by GPT-4o and Claude 3 Opus.
This report focuses on the training details of Flux.1, an open-source text-to-image generation model from Black Forest Labs, revealing the entire process from data preparation to deployment optimization.
LMSYS Org, in collaboration with MLCommons, has released a training benchmark report for the Llama 3.1 8B model. Based on MLCommons' standardized training benchmarks, the report details the entire pipeline from data processing to model convergence, providing reliable references for AI researchers and practitioners.
MLCommons, in collaboration with LMSYS Org, has launched the ISO-AUS benchmark, a novel AI model evaluation framework designed for isolation-aware serving scenarios. The benchmark focuses on model isolation, resource fairness, and low-latency response in multi-tenant environments.
MLCommons has released the MLPerf Training v5.1 benchmark results, the latest progress in AI model training performance evaluation. The submission covers nine core workloads and includes participants such as NVIDIA, Intel, AMD, and Google Cloud, demonstrating training capabilities from single-node to thousands-of-GPU large-scale clusters.
MLCommons announces the release of MLPerf Client 1.5, the latest benchmark suite for client inference scenarios, focusing on AI performance evaluation on mobile devices, laptops, and edge devices to provide more realistic testing standards.
MLCommons recently announced that its open-source privacy-preserving machine learning benchmarking platform, MedPerf, has officially added a WebUI feature. This update greatly improves usability, allowing developers to perform model evaluation and benchmarking through a browser without complex environment setup.
MLCommons released the latest VLM (Vision-Language Model) inference benchmark results, with the submission from the <strong>Shopify</strong> team attracting significant attention. Supported by LMSYS Org, the benchmark focuses on the inference performance of vision-language models in high-load e-commerce scenarios, aiming to provide standardized evaluation for production deployment.
KTransformers, developed by Tsinghua University's MadSys and Approaching.AI, optimizes CPU/GPU collaborative inference for sparse MoE models through AMX-optimized kernels, efficient device coordination, and expert deferral mechanisms, now integrated into SGLang for enhanced performance.
SGLang-Diffusion has achieved 2.5x performance improvements since its launch in November 2025, with support for new models, LoRA, parallel processing, and ComfyUI integration.
SGLang launches a highly optimized Pipeline Parallelism implementation designed for ultra-long context inference challenges. Through integrated optimizations and a clean design, it achieves a 3.31x speedup in prefill throughput for DeepSeek V3 on multi-node H20 clusters, demonstrating strong scalability for trillion-parameter models.
We developed Petit, a collection of FP16/BF16 × FP4 mixed-precision GPU kernels for AMD GPUs, enabling 1.74× faster Llama 3.3 70B inference on existing MI250/MI300 hardware without upgrades.
SGLang implements fully deterministic inference with only 34.35% performance overhead and enables 100% reproducible RL training in collaboration with slime, providing reliable solutions for rigorous scientific experiments.