Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(488) Artificial Intelligence(450) Anthropic(377) AI Safety(271) AI Agents(176) Meta(128) AI Ethics(123) WDCD(117) AI Regulation(114) Google(110) Generative AI(108) xAI(102) Data Centers(98) Smoke Test(96) Code Execution(95) Funding(90) Claude(90) AI Chips(90) AI(86) Material Constraints(86) Cybersecurity(82)

Whisper Inferencev5 1

MLCommons has released the Whisper Inference v5.1 benchmark, introducing the large-v3 model and providing standardized performance data for speech-to-text model inference, aiming to help developers optimize deployment.

MLC Whisper MLCommons
1,005 02-10

Small Llm Inference 5 1

MLCommons released the Small LLM Inference 5.1 benchmark, evaluating SLM performance in real-world scenarios such as chatbots and text generation. Key updates include the addition of Llama 3.2 1B, optimized test scenarios, and expanded hardware support.

MLC MLCommons 小型LLM
862 02-10

Deepseek Inference 5 1

The latest report from LMSYS Org shows that DeepSeek Inference 5.1 stands out in the MLCommons September 2025 inference benchmark. Designed for large language models, it focuses on low-latency and high-throughput inference optimization.

MLC DeepSeek 推理引擎
1,013 02-10

Mlperf Inference V5 1 Results

MLCommons announces the release of MLPerf Inference v5.1 benchmark results, featuring new generative AI workloads and expanded model coverage, with NVIDIA, AMD, and others competing on latency and throughput.

MLC MLPerf 推理基准
1,489 02-10

Mlperf Tiny V1 3 Results

The MLCommons organization has released the MLPerf Tiny v1.3 benchmark results, a key milestone for edge AI. It evaluates model performance on resource-constrained microcontrollers and embedded devices.

MLC MLPerf Tiny 边缘AI
1,104 02-10

Mlperf Tiny V1 3 Tech

MLPerf Tiny v1.3 is the latest edge AI benchmark version from MLCommons, designed for resource-constrained devices. It introduces two new benchmarks (Image Classification and Visual Wake Words) and optimizes existing ones, with evaluation rules covering closed and open divisions.

MLC MLPerf Tiny 边缘AI
944 02-10

Croissant Mcp

MLCommons has officially launched Croissant MCP, a major upgrade to the Croissant metadata format designed specifically for AI model cards. This standard, contributed by partners including LMSYS Org, aims to address the fragmentation of current model documentation.

MLC MLCommons Croissant MCP
863 02-10

Ailuminate Jailbreak V05

MLCommons and LMSYS Org have launched the AILuminate Jailbreak V05 benchmark, introducing complex multi-turn attack chains and roleplay prompts. Claude 3.5 Sonnet tops the ranking with 1485 Elo, followed closely by GPT-4o and Claude 3 Opus.

MLC AILuminate 越狱基准
1,131 02-10

Training Flux1

This report focuses on the training details of Flux.1, an open-source text-to-image generation model from Black Forest Labs, revealing the entire process from data preparation to deployment optimization.

MLC Flux.1 模型训练
1,146 02-10

Training Llama 3 1 8b

LMSYS Org, in collaboration with MLCommons, has released a training benchmark report for the Llama 3.1 8B model. Based on MLCommons' standardized training benchmarks, the report details the entire pipeline from data processing to model convergence, providing reliable references for AI researchers and practitioners.

MLC Llama 3.1 模型训练
1,146 02-10

Iso Aus

MLCommons, in collaboration with LMSYS Org, has launched the ISO-AUS benchmark, a novel AI model evaluation framework designed for isolation-aware serving scenarios. The benchmark focuses on model isolation, resource fairness, and low-latency response in multi-tenant environments.

MLC ISO-AUS AI基准
1,016 02-10

Training V5 1 Results

MLCommons has released the MLPerf Training v5.1 benchmark results, the latest progress in AI model training performance evaluation. The submission covers nine core workloads and includes participants such as NVIDIA, Intel, AMD, and Google Cloud, demonstrating training capabilities from single-node to thousands-of-GPU large-scale clusters.

MLC MLPerf 训练基准
937 02-10

Mlperf Client 1 5 Release

MLCommons announces the release of MLPerf Client 1.5, the latest benchmark suite for client inference scenarios, focusing on AI performance evaluation on mobile devices, laptops, and edge devices to provide more realistic testing standards.

MLC MLPerf 客户端基准
766 02-10

Medperf Adds Webui Capabilities

MLCommons recently announced that its open-source privacy-preserving machine learning benchmarking platform, MedPerf, has officially added a WebUI feature. This update greatly improves usability, allowing developers to perform model evaluation and benchmarking through a browser without complex environment setup.

MLC MedPerf WebUI
954 02-10

Vlm Inference Shopify

MLCommons released the latest VLM (Vision-Language Model) inference benchmark results, with the submission from the <strong>Shopify</strong> team attracting significant attention. Supported by LMSYS Org, the benchmark focuses on the inference performance of vision-language models in high-load e-commerce scenarios, aiming to provide standardized evaluation for production deployment.

MLC VLM推理 MLPerf基准
689 02-10

KTransformers Accelerates SGLang's Heterogeneous Inference

KTransformers, developed by Tsinghua University's MadSys and Approaching.AI, optimizes CPU/GPU collaborative inference for sparse MoE models through AMX-optimized kernels, efficient device coordination, and expert deferral mechanisms, now integrated into SGLang for enhanced performance.

LMSYS AI Technology 混合推理
1,656 02-04

SGLang-Diffusion: Two Months of Progress

SGLang-Diffusion has achieved 2.5x performance improvements since its launch in November 2025, with support for new models, LoRA, parallel processing, and ComfyUI integration.

LMSYS AI Technology 深度学习
1,218 02-04

SGLang Pipeline Parallelism: Million-Token Context Extension and Performance Breakthroughs

SGLang launches a highly optimized Pipeline Parallelism implementation designed for ultra-long context inference challenges. Through integrated optimizations and a clean design, it achieves a 3.31x speedup in prefill throughput for DeepSeek V3 on multi-node H20 clusters, demonstrating strong scalability for trillion-parameter models.

LMSYS SGLang Pipeline Parallelism
1,504 02-04

FP4 Mixed-Precision Inference Optimization on AMD GPUs

We developed Petit, a collection of FP16/BF16 × FP4 mixed-precision GPU kernels for AMD GPUs, enabling 1.74× faster Llama 3.3 70B inference on existing MI250/MI300 hardware without upgrades.

LMSYS AMD GPU FP4量化
1,070 02-04

SGLang Achieves Deterministic Inference and Reproducible RL Training

SGLang implements fully deterministic inference with only 34.35% performance overhead and enables 100% reproducible RL training in collaboration with slime, providing reliable solutions for rigorous scientific experiments.

LMSYS SGLang 确定性推理
1,259 02-04
20 21 22 23 24

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0