Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(522) Artificial Intelligence(467) Anthropic(397) AI Safety(310) AI Agents(183) Meta(135) AI Ethics(125) AI Regulation(125) WDCD(122) Google(119) Generative AI(111) xAI(105) Smoke Test(104) Data Centers(103) Code Execution(99) Cybersecurity(98) Funding(94) Claude(94) AI(92) AI Chips(91) Material Constraints(91)

R3 Integrity Rate Only 40.9%: Four Models Score Zero in WDCD Business Rule Scenario

In three rounds of testing on 8 v2 anchor questions, the average R3 integrity rate across 11 models was only 40.9%, with 4 models experiencing complete collapse (score 0).

WDCD Compliance Test 约束衰减
290 07-22

Grok 4 Scores 93.80 to Top the Compliance Test, Doubao Pro Trails at 67.30 with a 26.5-Point Gap

In the WDCD v3.1 compliance test, Grok 4 achieved the highest score of 93.80 among 11 evaluated models, while Doubao Pro scored the lowest at 67.30, a difference of 26.5 points. The top three models formed a clear tier with a significant gap from the rest.

WDCD Compliance Test AI模型评估
219 07-22

GLM-4.6 Integrity Rating Drops from Pass to Fail, Code Execution Surges by 47 Points

GLM-4.6's integrity rating fell from pass to fail in today's Smoke evaluation, while its code execution score surged by 47 points. However, the overall ranking increase was driven solely by this dimension, suggesting sampling fluctuation rather than genuine improvement.

GLM-4.6 Integrity Rating Smoke 评测
205 07-22

GPT-o3 Smoke Evaluation Main Leaderboard Plunges 8.3 Points, Code Execution Drops from 100 to 88.3

In today’s Smoke evaluation, GPT-o3’s main leaderboard score fell from 96.27 to 87.94, a drop of 8.3 points. Code execution declined from 100.00 to 88.30, while engineering judgment saw the steepest decline, dropping from 94.80 to 75.00.

GPT-o3 Code Execution Smoke Test
222 07-22

Grok 4 Leads with 98.35 Points: 2026-07-22 Smoke Quick Test Data Brief

On July 22, 2026, the YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first at 98.35 points. The Smoke test uses 10 daily questions to monitor short-term signals and is not equivalent to the full weekly ranking conclusion.

YZ Index Smoke快测 AI Evaluation
207 07-22

Claude Opus 4.7 Smoke Evaluation Main Ranking Drops 26.1 Points, Code Execution and Material Constraints Both Fail

In today's Smoke evaluation, Claude Opus 4.7's main ranking score dropped sharply by 26.1 points to 73.92. The code execution and material constraints dimensions saw significant declines, while engineering judgment remained relatively stable.

Claude Opus 4.7 Code Execution Smoke Test
258 07-21

Gemini 3.1 Pro Material Constraint Drops 17.8 Points, Main Ranking Falls 6 Points

In today's Smoke evaluation, Gemini 3.1 Pro's material constraint score dropped from 90.40 to 72.60, a decrease of 17.8 points, causing its main ranking to fall from 81.93 to 75.90.

Gemini 3.1 Pro Material Constraints Smoke Test
208 07-21

Claude Sonnet 4.6 and GPT-o3 Tie at 96.27: 2026-07-21 Smoke Quick Test Data Brief

On July 21, 2026, the YZ Index Smoke Quick Test covered 11 models, with Claude Sonnet 4.6 and GPT-o3 tying for first place at 96.27 points, showing balanced strengths in code execution and material constraints.

YZ Index Smoke快测 AI Evaluation
223 07-21

Qwen3 Max Main Score Plunges 14.9 Points, Code Execution Drops from 96.9 to 65.6

In today's Smoke evaluation, Qwen3 Max's main score dropped from 82.23 to 67.31, a decrease of 14.9 points, with the code execution dimension falling from 96.90 to 65.60.

Qwen3 Max Code Execution Smoke Test
219 07-20

Gemini 2.5 Pro Code Execution Dropped 24.6 Points in a Single Day; Overall Ranking Slid 6.5 Points

In today's Smoke evaluation, Gemini 2.5 Pro's code execution score dropped from 74.60 to 50.00 points (a decrease of 24.6 points), while its overall ranking fell from 76.49 to 69.98 points.

Gemini 2.5 Pro Code Execution Smoke Test
209 07-20

Claude Opus 4.7 Leads with 100 Points: 2026-07-20 Smoke Quick Test Data Brief

On July 20, 2026, the YZ Index Smoke Quick Test covered 11 models, with Claude Opus 4.7 scoring 100 points to top the daily rankings. This Smoke test uses 10 questions per day, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

YZ Index Smoke快测 AI Evaluation
211 07-20

Gemini 3.1 Pro Material Constraint Drops 26.6 Points, Main Ranking Still Up 5.4 Points

In today's Smoke evaluation, Gemini 3.1 Pro's material constraint score fell from 90.90 to 64.30, a drop of 26.6 points, while its main ranking overall still rose by 5.4 points.

Gemini 3.1 Pro Material Constraints Smoke Test
186 07-19

GPT-o3 Main Score Plummets 13.8 Points, Code Execution Drops from 70.3 to 48.5

GPT-o3's main score in today's Smoke evaluation fell from 80.61 to 66.86, with code execution dropping from 70.30 to 48.50, a single-day decline of 21.8 points.

GPT-o3 Code Execution Smoke Test
210 07-19

Claude Opus 4.7 Leads with Average Score of 86.9, GPT-o3 Drops 30.5 Points in 7 Days

The seven-day Smoke evaluation from July 13 to July 19, 2026 shows Claude Opus 4.7 leading with an average score of 86.9, while GPT-o3 declined sharply by 30.5 points. The report also analyzes rising models such as Qwen3 Max, declining models like Doubao Pro, along with volatility and integrity ratings.

Claude Opus 4.7 GPT-o3 Smoke 周趋势
157 07-19

Claude Opus 4.7 Tops with 95.19: 2026-07-19 Smoke Quick Test Data Brief

On 2026-07-19, the YZ Index Smoke Quick Test covered 10 models, with Claude Opus 4.7 ranking first at 95.19. Smoke is a daily 10-question quick test suitable for monitoring short-term signals and is not equivalent to Full weekly rankings.

YZ Index Smoke快测 AI Evaluation
266 07-19

GPT-o3 Tops with 80.61: 2026-07-18 Smoke Quick Test Data Brief

GPT-o3 leads with a score of 80.61 in the YZ Index Smoke quick test on July 18, 2026, covering 11 models. The daily 10-question test focuses on code execution and material constraints, providing short-term signals rather than long-term conclusions.

YZ Index Smoke快测 AI Evaluation
257 07-18

Grok 4 Smoke Evaluation Main Score Plunges 17.5 Points, Material Compliance Drops 21.9 in a Single Day

Grok 4's main score in today's Smoke evaluation dropped from 94.15 to 76.65, a decline of 17.5 points. The material compliance dimension fell by 21.9 points in a single day.

Grok 4 Material Constraints Smoke Test
200 07-17

DeepSeek V4 Pro Main Score Plummets 11.9 Points, Code Execution Drops 13.3

DeepSeek V4 Pro's main score in today's Smoke evaluation dropped 11.9 points from yesterday's 93.84 to 81.93. Code execution decreased by 13.3 points, and material constraints by 10.2 points, while engineering judgment and task expression remained unchanged.

DeepSeek V4 Pro Code Execution Smoke Test
211 07-17

Gemini 2.5 Pro and Gemini 3.1 Pro Tie at 92.44: 2026-07-17 Smoke Quick Test Data Brief

On July 17, 2026, the YZ Index Smoke quick test covered 11 models, with Gemini 2.5 Pro and Gemini 3.1 Pro tying for first place at 92.44 points. This Smoke test focused only on two main board dimensions: Code Execution and Material Constraints.

YZ Index Smoke快测 AI Evaluation
267 07-17

Doubao Pro Main Score Plunges 15 Points: Code Execution Drops from 75 to 58.3

Doubao Pro's main score fell from 86.25 to 71.22 in today's Smoke evaluation, with code execution dropping from 75.00 to 58.30 and material adherence dropping from 100.00 to 87.00. The analysis points to question selection volatility rather than model degradation.

Doubao Pro Code Execution Smoke Test
201 07-16
2 3 4 5 6

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0