Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(697) Artificial Intelligence(566) Anthropic(512) AI Safety(510) AI Agents(234) AI Regulation(200) Meta(176) WDCD(173) Smoke Test(157) Cybersecurity(152) Google(151) AI Ethics(149) Generative AI(140) Data Centers(138) Code Execution(134) Material Constraints(126) Funding(124) Claude(119) Compliance Test(117) AI Chips(115) xAI(114)

R3 Integrity Rate Only 50.6%: Grok 4 Zero Collapse, GPT-o3 and Qwen3 Max at 20% Collapse

In the WDCD v3.1 pilot, tests on eight v2 three-round anchor problems showed that 11 models achieved an average R3 integrity rate of just 50.6%. Grok 4 demonstrated a perfect resistance score of 1.63/2 with zero collapses, while GPT-o3 and Qwen3 Max each recorded a 20% collapse rate.

WDCD Compliance Test 约束衰减
706 07-26

DeepSeek V4 Pro Tops with 83.23: 2026-07-26 Smoke Quick Test Data Brief

On 2026-07-26, the YZ Index Smoke quick test covered 10 models, with DeepSeek V4 Pro ranking first with a score of 83.23. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusion.

YZ Index Smoke快测 AI Evaluation
537 07-26

Claude Sonnet 4.6 and Grok 4 Tie at 96.98: 2026-07-25 Smoke Test Data Brief

On July 25, 2026, the YZ Index Smoke test covered 11 models, with Claude Sonnet 4.6 and Grok 4 both scoring 96.98, tying for first place. This quick test monitors short-term signals and does not replace the full weekly ranking.

YZ Index Smoke快测 AI Evaluation
545 07-25

The Benchmark Behind the Next Wave of Ultra-Low-Power AI

Machine learning has moved beyond data centers into battery-powered devices with milliwatt power budgets. MLPerf Tiny provides a fair, architecture-neutral benchmark to compare performance and efficiency across radically different ultra-low-power systems.

MLC MLPerf Tiny TinyML
824 07-25

Agentic Inference for MLPerf Inference

MLPerf Inference introduces a new Agentic Inference benchmark targeting multi-turn agentic workloads such as coding assistants and workflow agents. It uses real-world traces to evaluate inference serving stacks under long context, KV-cache reuse, and variable output lengths.

MLC MLPerf 智能体推理
592 07-25

Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1

MLCommons introduces the new Edge Agentic Inference benchmark in MLPerf Inference v6.1, using Qwen3.6-27B with Q4_K_M quantization to measure accuracy and latency for on-device agentic LLM workloads. Submissions are due July 31, 2026.

MLC MLPerf 边缘AI
904 07-25

MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor Research

At Google Cloud Next 2026, MLCommons Medical AI Working Group and Google Cloud announced that MedPerf, MLCommons' federated benchmarking orchestrator, now supports Google Cloud's Confidential Computing capabilities, demonstrated through a clinically relevant brain tumor segmentation scenario to protect both patient data and AI model IP.

MLC 医疗AI 机密计算
814 07-25

Accelerating SGLang HiCache with Netpreme X-Mem™ MPU

Accelerating SGLang HiCache with Netpreme X-Mem™ MPUNetpreme TeamJuly 8, 2026 Netpreme X-Mem™ Memory Processing Unit (MPU) makes SGLang HiCache faster and more scalable by augmenting the slower Host D

LMSYS SGLang KV Cache
529 07-25

DSpark in SGLang: Speculative Decoding with Confidence-Driven, Variable-Length Verification

DSpark in SGLang: Speculative Decoding with Confidence-Driven, Variable-Length VerificationSGLang TeamJuly 6, 2026Speculative decoding trades extra compute for fewer decode steps, and the trade sours

LMSYS SGLang DSpark
787 07-25

Bringing DeepSeek-V4 Flash RL Training to AMD Instinct MI355X GPUs with Miles

Bringing DeepSeek-V4 Flash RL Training to AMD Instinct MI355X GPUs with MilesAMD & Miles TeamJuly 10, 2026DeepSeek-V4 RL is now supported in Miles on AMD Instinct™ MI355X GPUs with ROCm™! RL requi

LMSYS AMD ROCm
640 07-25

Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 Weeks

Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 WeeksSGLang TeamJuly 14, 2026TL;DR More than 500 TPS on 8xB300 (bs=1) Sync free speculative decoding for GLM 5.2 MTP Built-in I

LMSYS SGLang GLM-5.2
830 07-25

SGLang and Miles Add Day-0 Support for Inkling, a Frontier Multimodal Model

SGLang and Miles Add Day-0 Support for Inkling, a Frontier Multimodal ModelSGLang Team & Thinking Machines LabJuly 15, 2026We're excited to partner with the Thinking Machines team to bring Day-0 s

LMSYS SGLang Inkling
598 07-25

OPD Support in Miles

OPD Support in MilesKaixi Hou & Miles TeamJuly 18, 2026We recently implemented On-Policy Distillation (OPD) as an important feature in Miles. OPD is now integrated into Miles rollout and training

LMSYS Miles OPD
494 07-25

Grok 4 Leads with 84.21 Points: 2026-07-24 Smoke Quick Test Data Brief

On July 24, 2026, the Winzheng YZ Index Smoke Quick Test covered 10 models, with Grok 4 scoring 84.21 points to top the daily ranking. This daily 10-question test is designed for short-term signal observation and does not equate to the full weekly ranking conclusions.

YZ Index Smoke快测 AI Evaluation
439 07-24

GLM-4.6: 93.30 on Material Constraint but Integrity Fail, Code Execution 25.00 Drags Down Leaderboard

In Run#243 Smoke test, GLM-4.6 scored 55.74 on the main leaderboard, with code execution at 25.00, material constraint at 93.30, and an integrity rating of fail (probe score 30.00).

GLM-4.6 Integrity Rating Code Execution
478 07-23

Claude Opus 4.7 Tops with 96.99: 2026-07-23 Smoke Quick Test Data Brief

On 2026-07-23, the YZ Index Smoke Quick Test covered 11 models, with Claude Opus 4.7 ranking first at 96.99. Smoke is a daily 10-question quick test for short-term signals and does not replace the Full weekly ranking.

YZ Index Smoke快测 AI Evaluation
410 07-23

GLM-4.6 Soars 13.7 Points in WDCD; GPT-o3 Drops 6.9 – Commitment Top Restructured

In the latest WDCD v3.1 commitment test, GLM-4.6 surged 13.7 points over Run #233 to 92.00, while GPT-o3 fell 6.9 points to 87.10, directly reshuffling the top five rankings.

WDCD Compliance Test 模型评估
688 07-22

Resource Limitation Scenario Lowest at 1.55 Points: Maximum Spread of 2.45 Points Across 11 Models in WDCD Compliance Test

In the resource limitation scenario, gpt-5.5 scored only 1.55/4, and in business rules, Doubao-pro scored only 1.45/4, directly revealing the weakest constraint types in the WDCD v3.1 compliance test.

WDCD Compliance Test 模型横评
670 07-22

R3 Integrity Rate Only 40.9%: Four Models Score Zero in WDCD Business Rule Scenario

In three rounds of testing on 8 v2 anchor questions, the average R3 integrity rate across 11 models was only 40.9%, with 4 models experiencing complete collapse (score 0).

WDCD Compliance Test 约束衰减
689 07-22

Grok 4 Scores 93.80 to Top the Compliance Test, Doubao Pro Trails at 67.30 with a 26.5-Point Gap

In the WDCD v3.1 compliance test, Grok 4 achieved the highest score of 93.80 among 11 evaluated models, while Doubao Pro scored the lowest at 67.30, a difference of 26.5 points. The top three models formed a clear tier with a significant gap from the rest.

WDCD Compliance Test AI模型评估
555 07-22
10 11 12 13 14

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC Xinyuan luo

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0