Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(522) Artificial Intelligence(467) Anthropic(397) AI Safety(310) AI Agents(183) Meta(135) AI Ethics(125) AI Regulation(125) WDCD(122) Google(119) Generative AI(111) xAI(105) Smoke Test(104) Data Centers(103) Code Execution(99) Cybersecurity(98) Funding(94) Claude(94) AI(92) AI Chips(91) Material Constraints(91)

Claude Opus 4.7, Claude Sonnet 4.6, and GPT-o3 Tie at 81.44 Points: 2026-07-11 Smoke Quick Test Data Brief

On 2026-07-11, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7, Claude Sonnet 4.6, and GPT-o3 tying for first place at 81.44 points. This brief reports the daily snapshot results.

YZ Index Smoke快测 AI Evaluation
263 07-11

Claude Opus 4.7 Smoke Review Main Board Plunges 19.3 Points, Code Execution Drops 22 Points in a Single Day

In today's Smoke review, Claude Opus 4.7's main board score fell from 90.51 to 71.26, a drop of 19.3 points. Code execution dropped 22 points, while material constraint dropped 15.9 points.

Claude Opus 4.7 Code Execution Smoke Test
220 07-10

Grok 4 Main Score Plunges 8.4 Points, Material Constraint Drops 17.6 Points in a Single Day

Grok 4's main score in today's Smoke evaluation dropped 8.4 points from 87.66 to 79.30, with the Material Constraint dimension falling 17.6 points.

Grok 4 Material Constraints Smoke 评测
233 07-10

GPT-o3 Tops with 86.9 Points: 2026-07-10 Smoke Quick Test Data Brief

On 2026-07-10, the YZ Index Smoke Quick Test covered 9 models, with GPT-o3 ranking first at 86.9 points. Smoke is a daily 10-question quick test for monitoring short-term signals, not equivalent to Full weekly rankings.

YZ Index Smoke快测 AI Evaluation
300 07-10

GPT-o3 Material Constraint Drops 16.8 Points, Task Expression Falls 28.3 Points

GPT-o3's material constraint score dropped 16.8 points in today's Smoke evaluation, while task expression fell 28.3 points, causing the main ranking total to decline from 83.44 to 80.39.

GPT-o3 Material Constraints Smoke Test
250 07-09

Qwen3 Max Material Constraint Plunges 15.1 Points While Code Execution Surges 18.4 Points

In today's Smoke evaluation, Qwen3 Max's Material Constraint score dropped from 83.60 to 68.50, a decrease of 15.1 points, while its Code Execution score rose from 73.10 to 91.50, and its main ranking score increased from 77.83 to 81.15.

Qwen3 Max Material Constraints Smoke Test
211 07-09

Claude Opus 4.7 Tops with 90.51 Points: 2026-07-09 Smoke Quick Test Data Brief

On July 9, 2026, the YZ Index Smoke Quick Test covered 10 models, with Claude Opus 4.7 ranking first with a score of 90.51. The Smoke test is a daily 10-question quick assessment for monitoring short-term signals, not equivalent to the Full weekly ranking.

YZ Index Smoke快测 AI Evaluation
260 07-09

WDCD v3.1: DeepSeek V4 Pro Up 26.2 Points, Claude Sonnet 4.6 Down 5.9

Grok 4 WDCD scores 95.00, up 3.8 points from Run #211, maintaining first place; DeepSeek V4 Pro jumps 26.2 points to 94.00, and GLM-4.6 rises 21.8 points to 93.60, both within 2 points of Grok 4. Only Claude Sonnet 4.6 declines, by 5.9 points.

WDCD Compliance Test 模型评估
344 07-08

WDCD v3.1 Five-Scenario Cross-Evaluation: Business Rules Score 1.3 at the Bottom, 11 Models Show Subject Imbalance of 2.1 Points

In the WDCD v3.1 pilot, the Business Rules scenario scored the lowest overall, with champion claude-opus-4.7 achieving only 3.5/4 and bottom-ranked qwen3-max scoring just 1.3/4, far below the champion scores of the other four scenarios.

WDCD Compliance Test 业务规则场景
342 07-08

R3 Integrity Rate Only 61.4%: Claude Sonnet's 20% Collapse Rate Exposes Three-Round Degradation Fault

In a worst-of-3 sampling of only 8 v2 anchor questions, the average R3 integrity rate across 11 models was merely 61.4%, while R1 confirmation rate remained as high as 95% and R2 resistance rate 73%, revealing the true performance of mainstream models under hard constraints.

WDCD Compliance Test 模型衰减
285 07-08

Grok 4 Tops WDCD Compliance Leaderboard with 95 Points, Claude Sonnet 4.6 Trails at 64.1 Points

Grok 4 leads the WDCD Compliance Leaderboard with 95.00 points, while Claude Sonnet 4.6 ranks 11th with 64.10 points, a gap of 30.9 points.

WDCD 守约测试 守约测试排行榜 Grok 4
211 07-08

DeepSeek V4 Pro Leads with 95.19 Points: 2026-07-08 Smoke Quick Test Data Brief

The 2026-07-08 YZ Index Smoke Quick Test covered 10 models, with DeepSeek V4 Pro ranking first at 95.19 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.

YZ Index Smoke快测 AI Evaluation
365 07-08

Claude Opus 4.7 and Grok 4 Tie at 96.99: 2026-07-07 Smoke Quick Test Data Brief

On 2026-07-07, the Winzheng YZ Index Smoke Quick Test covered 11 models. Claude Opus 4.7 and Grok 4 tied for first place with a score of 96.99.

YZ Index Smoke快测 AI Evaluation
298 07-07

Doubao Pro Leads with 83.91 Points: 2026-07-06 Smoke Quick Test Data Brief

In the YZ Index Smoke Quick Test on July 6, 2026, Doubao Pro ranked first with a Main Board score of 83.91, covering 11 models in 10 daily questions. The test focuses on code execution and material constraints, serving as a short-term monitoring signal rather than a long-term conclusion.

YZ Index Smoke快测 AI Evaluation
800 07-06

GLM-4.6 Scores 25 in Material Constraint, 88.7 in Code Execution, Zero on Integrity Probe

In the Smoke Quick Test Run#214 on 2026-07-05, GLM-4.6 scored 60.04 on the main leaderboard, with code execution at 88.70, material constraint at 25.00, integrity rating fail, and probe score 0.00.

GLM-4.6 Material Constraints Integrity Rating
350 07-05

Doubao Pro and Gemini 3.1 Pro tied at 88.54: 2026-07-05 Smoke Quick Test Data Brief

On July 5, 2026, the YZ Index Smoke Quick Test covered 11 models, with Doubao Pro and Gemini 3.1 Pro tying for first place at 88.54 points. Smoke is a daily 10-question quick test for observing short-term signals and is not equivalent to the Full weekly ranking.

YZ Index Smoke快测 AI Evaluation
673 07-05

Agent-Assisted SGLang Development: An Initial Exploration

Agent-Assisted SGLang Development: An Initial ExplorationSGLang TeamJuly 2, 2026SGLang development increasingly goes beyond isolated code changes. The same repository now spans LLM serving,…

LMSYS SGLang Agent开发
402 07-04

Qwen3 Max Main Leaderboard Plummets 12.9 Points, Code Execution Drops 26.8 in a Single Day

In the June 2026 Smoke evaluation of the YZ Index, Qwen3 Max's main leaderboard score fell from 84.92 to 72.02, a drop of 12.9 points, with the code execution dimension plummeting from 96.30 to 69.50.

Qwen3 Max Code Execution Smoke Test
328 07-04

Qwen3 Max Main Board Plunges 12.9 Points, Gemini 2.5 Pro Leads Smoke Lite List with 96.99 Points

In the Smoke Lite evaluation of 11 models on July 4, 2026, by the YZ Index, Gemini 2.5 Pro ranked first with a Main Board score of 96.99, while Qwen3 Max's Main Board score plunged 12.9 points to 72.02.

Gemini 2.5 Pro Qwen3 Max Smoke Test
327 07-04

WDCD Review: Business Rules Scenario Lowest at 1.55, grok-4 Wins Security Compliance with 3.86

In the WDCD v3.1 compliance test, the business rules scenario scored the lowest among all models, with grok-4 leading at 3.5/4, while doubao-pro and qwen3-max only scored 1.55/4.

WDCD Compliance Test 业务规则
348 07-03
4 5 6 7 8

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0