Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(697) Artificial Intelligence(566) Anthropic(512) AI Safety(509) AI Agents(234) AI Regulation(200) Meta(176) WDCD(173) Smoke Test(157) Cybersecurity(152) Google(151) AI Ethics(149) Generative AI(140) Data Centers(137) Code Execution(134) Material Constraints(126) Funding(123) Claude(119) Compliance Test(117) AI Chips(115) xAI(114)

Claude Sonnet 4.6 and DeepSeek V4 Pro Tie at 92.17: 2026-08-06 Smoke Quick Test Data Brief

The 2026-08-06 YZ Index Smoke quick test covered 9 models, with Claude Sonnet 4.6 and DeepSeek V4 Pro tying for first place at 92.17 points. Doubao Pro and GLM-4.6 were not ranked due to incomplete data from API failures or timeouts.

YZ Index Smoke快测 AI Evaluation
390 08-06

GPT-o3 Stages Comeback with 9.5-Point Rise, GLM-4.6 Plunges 14.9 — Five Models Reshuffled on WDCD Compliance Leaderboard

This round of WDCD v3.1 testing shows GPT-o3 rising 9.5 points and Gemini 2.5 Pro rising 7.6 points, while GLM-4.6 plunges 14.9 points, Claude Sonnet 4.6 drops 10.8 points, and Claude Opus 4.7 drops 5.9 points.

WDCD Compliance Test Claude模型
1,146 08-05

WDCD Comparative Review: Safety Compliance Lowest at 1.8 Points, Engineering Standards Full 4 Across the Board

The WDCD v3.1 compliance test shows safety compliance as the hardest scenario, with gpt-5.5 and qwen3-max scoring only 1.8/4, while all 11 models in the engineering standards scenario scored at least 3.2/4. The results reveal a capability ceiling in engineering standards and a persistent gap in safety compliance.

WDCD Compliance Test 模型偏科
1,073 08-05

WDCD Three-Round Attrition: R3 Integrity Rate Only 54.5%, Doubao Pro Collapses at R1, Six Models Zero Collapse

Under a sampling scope targeting only 8 v2 anchor questions, 11 models posted an average R1 confirmation rate of 0.91, an average R2 resistance rate that fell to 0.68, and an average R3 integrity rate of just 54.5%. This trajectory reveals systematic attrition of constraints under sustained pressure.

WDCD Compliance Test 约束衰减
1,055 08-05

Grok 4 Tops WDCD Commitment Ranking with 97.5 Points; Doubao Pro Trails at 68 Points

In the WDCD v3.1 commitment test, Grok 4 ranked first with 97.50 points while Doubao Pro ranked last with 68.00 points, a 29.5-point gap between the top and bottom.

WDCD Compliance Test Grok 4
1,385 08-05

DeepSeek V4 Pro, GPT-5.5, and GPT-o3 Tie at 80.52 Points: 2026-08-05 Smoke Quick Test Data Briefing

On 2026-08-05, the YZ Index Smoke quick test covered 9 models, with DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tying for first place at 80.52 points. The briefing highlights notable single-day declines for Qwen3 Max, Grok 4, and Gemini 3.1 Pro, pending verification in subsequent runs.

YZ Index Smoke快测 AI Evaluation
847 08-05

GLM-4.6 Smoke Evaluation: Material Constraint 51.80, Task Expression 50.00; API Failure Leaves Two Dimensions Missing

GLM-4.6's Smoke evaluation today recorded 51.80 for material constraint and 50.00 for task expression, while data for the execution and judgment dimensions is missing due to an API failure. An automatic re-run has been initiated, and this round will not be included in the main leaderboard ranking.

GLM-4.6 Material Constraints Smoke Test
1,067 08-04

Doubao Pro Smoke Evaluation Shows All 5 Dimensions Missing in 0-Score Anomaly, with API Timeout as Main Cause

Doubao Pro's Smoke evaluation today returned no scores across all five dimensions—execution, grounding, judgment, integrity, and communication—due to API timeouts, resulting in removal from the main leaderboard. The anomaly points to interface-level failure rather than model capability degradation.

Doubao Pro API故障 Smoke Test
1,135 08-04

Gemini 2.5 Pro Leads with 89.56: 2026-08-04 Smoke Quick-Test Data Brief

On 2026-08-04, the YZ Index Smoke quick-test covered 9 models, with Gemini 2.5 Pro topping the daily ranking at 89.56 points. Smoke is a daily 10-question quick test designed for short-term signal monitoring and does not carry the same weight as the Full weekly leaderboard.

YZ Index Smoke快测 AI Evaluation
1,068 08-04

Claude Opus 4.7 Leads at 95.19 Points: 2026-08-03 Smoke Quick Test Data Briefing

On 2026-08-03, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first at 95.19 points. Smoke is a daily 10-question quick test for observing short-term signals, not equivalent to the Full weekly ranking conclusions.

YZ Index Smoke快测 AI Evaluation
1,532 08-03

GLM-4.6 Smoke Evaluation: Main Leaderboard Score 74, Code Execution 82.3, Material Constraint 95, API Failure Leaves Dimensions Missing

GLM-4.6 scored 74.00 on the main leaderboard in today's Smoke evaluation, with 82.30 on code execution and 95.00 on material constraint. Two dimensions are missing due to API failure/timeout and have entered automatic retesting, excluded from this period's ranking.

GLM-4.6 Code Execution Smoke Test
713 08-02

Qwen3 Max Rallies +36.8 to Lead; Gemini 3.1 Pro Slips 5.6 as Biggest Loser

In the July 28–August 2, 2026 Smoke evaluation, Qwen3 Max posted the largest seven-day gain (+36.8) to close at 96.1, while Gemini 3.1 Pro fell 5.6 points to 94.45, making it the biggest loser. The report also analyzes scoring trajectories, volatility drivers, integrity-rating changes, and implications for users.

Qwen3 Max Gemini 3.1 Pro 模型趋势分析
552 08-02

Doubao Pro Leads with 96.7 Points: 2026-08-02 Smoke Quick Test Data Briefing

On August 2, 2026, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first at 96.7 points. The briefing covers daily scores across code execution and material constraint dimensions, along with key fluctuations and integrity signals to monitor.

YZ Index Smoke快测 AI Evaluation
703 08-02

MLPerf Endpoints v0.7: A Foundation Release

MLPerf Endpoints v0.7 marks the foundational release of a buyer-centric AI inference benchmark, publishing initial results from Coreweave, Google, Intel, KRAI, and Nvidia across three benchmarks. The release supports automated submission pipelines, continuous review tooling, and dynamic result visualization, with v1.0 planned for later this year.

MLC MLPerf AI基准测试
974 08-02

SGLang and Miles Add Day-0 Support for Kimi K3

SGLang and Miles Add Day-0 Support for Kimi K3SGLang TeamJuly 27, 2026We are excited to announce Day-0 support for Kimi K3 in SGLang and Miles. K3 is the first open-source model in the 3-trillion-para

LMSYS SGLang Kimi K3
1,079 08-02

Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles

Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in MilesZiang Li, humans& and Miles TeamJuly 29, 2026 TL;DR: We implemented two Blackwell-native RL recipes in Miles: end

LMSYS Blackwell MXFP8
689 08-02

RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUs

RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUsRadixArk & GoogleJuly 30, 2026RadixArk and Google Cloud are partnering to bring SGLang to TPUs, giving developers ultimate fl

LMSYS SGLang TPU
804 08-02

Toward a Cleaner Quantization Stack in SGLang

Toward a Cleaner Quantization Stack in SGLangSGLang X Ascend TeamJuly 28, 2026Quantization has moved from an advanced feature to an essential part of high-throughput LLM serving. As the number of chec

LMSYS SGLang 量化优化
588 08-02

GLM-4.6 Material Constraint Score Plummets 27.3 Points, Main Score Rises 30.2 Points

In today's Smoke evaluation, GLM-4.6's material constraint score dropped from 75.00 to 47.70 points, while its main score rose from 46.29 to 76.47 points.

GLM-4.6 Material Constraints Smoke Test
633 08-01

GPT-o3 Drops 13.9 Points on Today's Main Leaderboard, Losing Ground in Both Code Execution and Material Constraints

GPT-o3 scored 79.28 points on today's Smoke evaluation main leaderboard, down 13.9 points from yesterday's 93.16, with notable declines in both code execution and material constraint dimensions.

GPT-o3 Code Execution Smoke Test
608 08-01
8 9 10 11 12

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC Xinyuan luo

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0