Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(522) Artificial Intelligence(467) Anthropic(397) AI Safety(310) AI Agents(183) Meta(135) AI Ethics(125) AI Regulation(125) WDCD(122) Google(119) Generative AI(111) xAI(105) Smoke Test(104) Data Centers(103) Code Execution(99) Cybersecurity(98) Funding(94) Claude(94) AI(92) AI Chips(91) Material Constraints(91)

R3 Integrity Rate Only 30.2%: 11 Models, 3-Round Anchor Questions, 44 Complete Collapses

In 275 samples on 8 v2 anchor questions, the average R1 confirmation rate was 0.99, but the R3 integrity rate was only 30.2%, with 44 complete collapses (score 0). This data directly reveals the rapid degradation pattern of models after initial commitment as rounds increase.

WDCD Compliance Test v3.1约束衰减
348 07-03

Grok 4 Scores 91.20 to Top WDCD Compliance Rankings, Qwen3 Max Trails at 57.48 with 33.72-Point Gap

Grok 4 tops the WDCD Compliance Leaderboard with 91.20 points, while Qwen3 Max ranks last with 57.48 points, a gap of 33.72 points between the top and bottom.

WDCD Compliance Test 模型守约能力
572 07-03

GPT-5.5 Leads Smoke Benchmark with Perfect Execution Score of 86.95, Exposing Constraint Weakness

In the Smoke lightweight benchmark on July 3, 2026, GPT-5.5 ranked first with a main score of 86.95, driven by a perfect code execution score of 100, while its material constraint score of 71 highlights a common weakness.

GPT-5.5 Code Execution Smoke 轻量评测
283 07-03

Gemini 3.1 Pro Tops with 82.97 Points, Execution Score of 75 Points Widens Gap with Second Place

In the YZ Index Smoke lightweight evaluation on July 2, 2026, Gemini 3.1 Pro achieved first place on the main leaderboard with 82.97 points (Execution 75, Material Constraint 92.7), while Doubao Pro ranked second with 81.98 points (Execution 75, Material Constraint 90.5), both tied for the highest execution score.

Gemini 3.1 Pro Code Execution Material Constraints
550 07-02

WDCD Three-Round Test: Grok 4 Zero Crashes, GPT-5.5 Five R3 Collapses

In the WDCD three-round test, Grok 4 maintained a perfect score of 2 in all 10 R3 questions, while GPT-5.5 suffered 5 zero-score crashes, with an average R3 integrity score of only 1.00/2.

WDCD Compliance Test 模型衰减
808 07-01

Grok 4 Scores Perfect 100 to Dominate WDCD Commitment Ranking, GPT-5.5 Trails with Only 62.5 Points

In the latest WDCD commitment test, Grok 4 achieved a perfect 100 points, while GPT-5.5 ranked last at 62.5 points. The results reveal a clear hierarchy, with top models excelling across all phases and bottom models collapsing under interference and pressure.

WDCD Compliance Test 模型排行榜
874 07-01

Doubao Pro Smoke Evaluation Main Ranking Plunges 18.6 Points, Code Execution Drops 38.8 in a Single Day

In the YZ Index June 2026 live test of 11 models, Doubao Pro’s Smoke Evaluation main ranking fell from 85.91 yesterday to 67.32 today, a drop of 18.6 points, primarily due to the code execution dimension falling from 83.30 to 44.50.

Doubao Pro Code Execution Smoke快测
989 07-01

Grok 4 Smoke Evaluation Main Score Plummets 15.3 Points, Code Execution Drops 31.4 in a Single Day

In today's YZ Index Smoke evaluation, Grok 4's main score dropped from 97.98 to 82.73, a decrease of 15.3 points, and code execution fell from 100.00 to 68.60. The single-day volatility is significant but consistent with small-sample draw characteristics, not necessarily indicating model degradation.

Grok 4 Code Execution 单日波动
332 07-01

Claude Opus 4.7 Tops with 94.82 Points, Gemini 3.1 Pro Plunges 32.2 Points

In the Smoke lightweight evaluation on July 1, 2026, Claude Opus 4.7 ranked first on the main leaderboard with a score of 94.82, while Gemini 3.1 Pro experienced a sharp drop of 32.2 points. The evaluation highlights a clear divergence between constraint and execution scores across models.

Claude Opus Code Execution 模型排名
524 07-01

Claude Sonnet 4.6 Smoke Main Ranking Plunges 15.3 Points, Code Execution Drops 25 Points in a Single Day

In the June 2026 Smoke evaluation of the YZ Index, Claude Sonnet 4.6 saw its main ranking score drop from 97.84 to 82.52 points, a single-day decline of 15.3 points, driven primarily by a 25-point fall in the code execution dimension.

Claude Sonnet 4.6 Code Execution Smoke Test
293 06-30

Claude Opus 4.7 Main Score Plunges 16 Points in Smoke Test, Code Execution Drops 27.2 in a Single Day

In the YZ Index June 2026 Smoke Evaluation, Claude Opus 4.7's main score dropped from 100.00 yesterday to 84.01 today, and its code execution dimension fell from 100.00 to 72.80.

Claude Opus 4.7 Code Execution Smoke Test
316 06-30

Gemini 3.1 Pro Tops with 98.47 Points, Claude's Execution Score Plunges 27.2 to 72.8

In the June 30, 2026 Smoke Lite evaluation of the YZ Index, Gemini 3.1 Pro ranked first with a main score of 98.47 points. Multiple models saw significant drops in execution scores, with Claude's execution scores plummeting over 25 points.

Gemini 3.1 Pro Code Execution Smoke 轻量评测
316 06-30

The patch model is breaking. AI evaluation needs a new way to disclose what it finds.

The security community has long relied on Coordinated Vulnerability Disclosure for handling dangerous findings, but this model fails for AI systems. MLCommons recognizes this as a core governance challenge for frontier model evaluation and is working to establish responsible disclosure standards.

MLC AI Safety 模型评估
362 06-29

Chakra Comes of Age: A Standardized Trace Ecosystem for AI Systems Benchmarking and Co-design

The Chakra initiative by MLCommons aims to standardize AI system benchmarking through open execution traces, addressing fragmentation and enabling collaborative co-design across industry and academia.

MLC AI基准测试 Chakra
306 06-29

MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs

MLCommons today announced the launch of MLPerf Mobile v6.0, adding generative AI benchmarks for running large language models (LLMs) on Android devices. These tests join existing benchmarks for image generation, object detection, and super resolution in the MLPerf Mobile app to form a complete test suite.

MLC MLPerf Mobile 设备端 LLM
425 06-29

MLCommons Releases MLPerf Training v6.0 Results

MLPerf Training v6.0 introduces two new Mixture-of-Experts benchmarks, DeepSeek V3 and GPT-OSS 20B, reflecting the industry's shift toward sparse computing. With 95 unique systems submitted by 24 organizations, the results show increasing diversity in hardware and software.

MLC MLPerf 基准测试
344 06-29

Optimizing GLM4-MoE for Production: 65% Faster TTFT with SGLang

Optimizing GLM4-MoE for Production: 65% Faster TTFT with SGLangNovita AIJanuary 21, 2026TL;DR A suite of production-tested, high-impact optimizations has been developed by Novita AI for deploying…

LMSYS SGLang GLM4-MoE
291 06-29

Squeezing 1TB Model Rollout into a Single H200: INT4 QAT RL End-to-End Practice

Squeezing 1TB Model Rollout into a Single H200: INT4 QAT RL End-to-End PracticeSGLang RL Team, InfiXAI Team, Ant Group Asystem & AQ Infra Team, slime Team, RadixArk TeamJanuary 26, 2026 💡 TL;DR:…

LMSYS INT4 QAT SGLang RL
263 06-29

No Token Left Behind: Demystifying Token-In-Token-Out in Miles

No Token Left Behind: Demystifying Token-In-Token-Out in MilesMiles Team: Jiajun Li, Yuzhen Zhou, Shi Dong, Yanbin Jiang, Mao Cheng, Yusheng Su, Yueming Yuan, Zhichen Zeng, Banghua ZhuJune 5, 2026In…

LMSYS 强化学习 Token处理
486 06-29

Win on TCO: How AMD Instinct™ MI355X Achieves Cost-Competitive Distributed Inference Through SGLang with MoRI

Win on TCO: How AMD Instinct™ MI355X Achieves Cost-Competitive Distributed Inference Through SGLang with MoRIAMD & SGLang TeamMay 28, 2026The SGLang and AMD team has worked closely to unlock…

LMSYS AMD MI355X SGLang
302 06-29
5 6 7 8 9

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0