Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(670) Artificial Intelligence(559) Anthropic(486) AI Safety(458) AI Agents(225) AI Regulation(174) Meta(170) WDCD(163) Smoke Test(150) Cybersecurity(146) AI Ethics(146) Google(145) Generative AI(131) Code Execution(131) Data Centers(129) Material Constraints(122) Funding(117) Claude(116) xAI(113) AI Chips(111) Compliance Test(109)

WDCD Compliance Leaderboard: Grok 4 Wins with 91.04 Points, Doubao Pro Trails at 58—a 33-Point Gap

In the WDCD v3.1 compliance test, Grok 4 ranked first with 91.04 points while Doubao Pro placed last with 58.00 points, a 33.04-point gap between the top and bottom.

WDCD Compliance Test AI模型排行
337 08-09

Qwen3 Max Main Leaderboard Plunges 10.2 Points in a Day, Driven by 21.5-Point Drop in Material Constraint

Qwen3 Max's main leaderboard score in the Smoke evaluation fell from 96.04 to 85.82 in a single day, a decline of 10.2 points. The material constraint dimension accounted for most of the drop, falling 21.5 points, while code execution remained nearly unchanged.

Qwen3 Max Material Constraints Smoke Test
289 08-09

DeepSeek V4 Pro Material Constraint Score Drops 19.5 Points in a Single Day, Main Ranking Slips 7.6 Points

DeepSeek V4 Pro's material constraint score fell from 89.50 to 70.00 in today's Smoke evaluation, dragging the main ranking down from 94.07 to 86.50. The drop is likely attributable to sampling variance rather than systematic model degradation.

DeepSeek V4 Pro Material Constraints Smoke Test
294 08-09

Claude Sonnet 4.6 Drops 22.6 Points, GLM-4.6 Fluctuates 59.9, GPT-o3 Rises 6.7 Points — Smoke Weekly Trend

During August 6–9, 2026, Claude Sonnet 4.6 posted the largest single-model decline in the Smoke rapid test, falling 22.6 points, while GPT-o3 rose 6.7 points. GLM-4.6 recorded extreme volatility of 59.9.

Claude Sonnet 4.6 GPT-o3 Model Fluctuations
294 08-09

GPT-o3 Tops with 95.91 Points: 2026-08-09 Smoke Quick-Test Data Brief

On 2026-08-09, the YZ Index Smoke quick test covered 9 models, with GPT-o3 ranking first at 95.91 points. The test focuses on code execution and material constraint, with several top models showing notable declines.

YZ Index Smoke快测 AI Evaluation
333 08-09

Doubao Pro Smoke Evaluation: All Five Dimensions Absent, API Failure Yields Zero Records

Doubao Pro scored "-" on all five dimensions—execution, grounding, judgment, integrity, and communication—in today's Smoke evaluation, as well as on the main leaderboard. The model completed no tasks due to API failure/timeout and is excluded from this period's ranking.

Doubao Pro Smoke 评测 API 故障
331 08-08

Claude Sonnet 4.6 Code Execution Plunges 19.5 Points While Leaderboard Score Rises 13.8 Points

In today's Smoke evaluation, Claude Sonnet 4.6's code execution score dropped from 94.50 to 75.00, a decline of 19.5 points, while its material constraint score surged from 43.30 to 97.80, lifting the overall leaderboard score from 71.46 to 85.26.

Claude Sonnet 4.6 Code Execution Smoke Test
419 08-08

Claude Opus 4.7 and GPT-o3 Tie at 97.66: 2026-08-08 Smoke Quick Test Data Brief

On 2026-08-08, the YZ Index Smoke quick test covered 9 models, with Claude Opus 4.7 and GPT-o3 tying at 97.66 points for the top spot. Smoke is a daily 10-question quick test for observing short-term signals and does not carry the weight of the Full weekly leaderboard.

YZ Index Smoke快测 AI Evaluation
374 08-08

SpecForge v0.3.0: a Unified Disaggregated and Colocated Speculative Decoding Stack, and New Open SpecBundle Draft Models

SpecForge v0.3.0: a Unified Disaggregated and Colocated Speculative Decoding Stack, and New Open SpecBundle Draft ModelsThe SpecForge TeamAugust 4, 2026When we first released SpecForge, a training job

LMSYS SpecForge 推测解码
412 08-08

Full-Stack Performance Optimization of AR+DiT in SGL-Diffusion

Full-Stack Performance Optimization of AR+DiT in SGL-DiffusionAscend TeamAugust 05, 2026TL;DR Replaces the HF backend with SRT to accelerate AR modeling and resolve parallelism conflicts, with dedicat

LMSYS SGLang AR+DiT
322 08-08

HPC-Ops × SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan

HPC-Ops × SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent HunyuanTencent Hunyuan AI Infra and the SGLang TeamAugust 7, 2026HPC-Ops is an open-source operator library for

LMSYS AI推理优化 MoE算子
433 08-08

Claude Opus 4.7 Code Execution Plunges 30.5 Points, Main Leaderboard Drops Only 6.4 Points

Claude Opus 4.7's code execution score fell from 100.00 to 69.50 in today's Smoke evaluation, while the main leaderboard dropped only 6.4 points. The opposing movement of dimensions points to random question sampling rather than model degradation.

Claude Opus 4.7 Code Execution Smoke评测波动
344 08-07

Grok 4 Material Constraint Drops 17.6 Points; Main Leaderboard Falls Just 1.8 Points

Grok 4's material constraint score fell from 82.60 to 65.00 in today's Smoke evaluation, while its overall main leaderboard score only declined from 82.99 to 81.23.

Grok 4 Material Constraints Smoke Test
341 08-07

Gemini 2.5 Pro Leads at 87.21: 2026-08-07 Smoke Quick Test Data Briefing

The 2026-08-07 YZ Index Smoke quick test covered 9 models, with Gemini 2.5 Pro ranking first at 87.21 points. As a small-sample daily signal, the results are better suited for short-term monitoring than long-term conclusions.

YZ Index Smoke快测 AI Evaluation
393 08-07

GLM-4.6 Smoke Test: Material Constraint Scores 71.90, Code Execution and Integrity Dimensions Missing

GLM-4.6 scored 71.90 on material constraint in today's Smoke evaluation but missed both the code execution and integrity dimensions due to API failures or timeouts, excluding it from main leaderboard ranking.

GLM-4.6 Material Constraints Smoke Test
320 08-06

Doubao Pro Smoke Evaluation Shows Complete Data Loss Across All Dimensions; API Outage Excludes It from Main Leaderboard This Cycle

Doubao Pro's Smoke evaluation this cycle showed complete data loss across all five dimensions — execution, grounding, judgment, integrity, and communication — caused by an API failure or timeout. The model will not be included in the main leaderboard ranking this cycle.

Doubao Pro Smoke 评测 API 故障
350 08-06

Claude Sonnet 4.6 and DeepSeek V4 Pro Tie at 92.17: 2026-08-06 Smoke Quick Test Data Brief

The 2026-08-06 YZ Index Smoke quick test covered 9 models, with Claude Sonnet 4.6 and DeepSeek V4 Pro tying for first place at 92.17 points. Doubao Pro and GLM-4.6 were not ranked due to incomplete data from API failures or timeouts.

YZ Index Smoke快测 AI Evaluation
366 08-06

GPT-o3 Stages Comeback with 9.5-Point Rise, GLM-4.6 Plunges 14.9 — Five Models Reshuffled on WDCD Compliance Leaderboard

This round of WDCD v3.1 testing shows GPT-o3 rising 9.5 points and Gemini 2.5 Pro rising 7.6 points, while GLM-4.6 plunges 14.9 points, Claude Sonnet 4.6 drops 10.8 points, and Claude Opus 4.7 drops 5.9 points.

WDCD Compliance Test Claude模型
1,104 08-05

WDCD Comparative Review: Safety Compliance Lowest at 1.8 Points, Engineering Standards Full 4 Across the Board

The WDCD v3.1 compliance test shows safety compliance as the hardest scenario, with gpt-5.5 and qwen3-max scoring only 1.8/4, while all 11 models in the engineering standards scenario scored at least 3.2/4. The results reveal a capability ceiling in engineering standards and a persistent gap in safety compliance.

WDCD Compliance Test 模型偏科
1,038 08-05

WDCD Three-Round Attrition: R3 Integrity Rate Only 54.5%, Doubao Pro Collapses at R1, Six Models Zero Collapse

Under a sampling scope targeting only 8 v2 anchor questions, 11 models posted an average R1 confirmation rate of 0.91, an average R2 resistance rate that fell to 0.68, and an average R3 integrity rate of just 54.5%. This trajectory reveals systematic attrition of constraints under sustained pressure.

WDCD Compliance Test 约束衰减
1,012 08-05
6 7 8 9 10

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0