Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(670) Artificial Intelligence(559) Anthropic(486) AI Safety(458) AI Agents(225) AI Regulation(174) Meta(170) WDCD(163) Smoke Test(150) Cybersecurity(146) AI Ethics(146) Google(145) Generative AI(131) Code Execution(131) Data Centers(129) Material Constraints(122) Funding(117) Claude(116) xAI(113) AI Chips(111) Compliance Test(109)

Grok 4 Tops WDCD Commitment Ranking with 97.5 Points; Doubao Pro Trails at 68 Points

In the WDCD v3.1 commitment test, Grok 4 ranked first with 97.50 points while Doubao Pro ranked last with 68.00 points, a 29.5-point gap between the top and bottom.

WDCD Compliance Test Grok 4
1,243 08-05

DeepSeek V4 Pro, GPT-5.5, and GPT-o3 Tie at 80.52 Points: 2026-08-05 Smoke Quick Test Data Briefing

On 2026-08-05, the YZ Index Smoke quick test covered 9 models, with DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tying for first place at 80.52 points. The briefing highlights notable single-day declines for Qwen3 Max, Grok 4, and Gemini 3.1 Pro, pending verification in subsequent runs.

YZ Index Smoke快测 AI Evaluation
789 08-05

GLM-4.6 Smoke Evaluation: Material Constraint 51.80, Task Expression 50.00; API Failure Leaves Two Dimensions Missing

GLM-4.6's Smoke evaluation today recorded 51.80 for material constraint and 50.00 for task expression, while data for the execution and judgment dimensions is missing due to an API failure. An automatic re-run has been initiated, and this round will not be included in the main leaderboard ranking.

GLM-4.6 Material Constraints Smoke Test
1,017 08-04

Doubao Pro Smoke Evaluation Shows All 5 Dimensions Missing in 0-Score Anomaly, with API Timeout as Main Cause

Doubao Pro's Smoke evaluation today returned no scores across all five dimensions—execution, grounding, judgment, integrity, and communication—due to API timeouts, resulting in removal from the main leaderboard. The anomaly points to interface-level failure rather than model capability degradation.

Doubao Pro API故障 Smoke Test
1,076 08-04

Gemini 2.5 Pro Leads with 89.56: 2026-08-04 Smoke Quick-Test Data Brief

On 2026-08-04, the YZ Index Smoke quick-test covered 9 models, with Gemini 2.5 Pro topping the daily ranking at 89.56 points. Smoke is a daily 10-question quick test designed for short-term signal monitoring and does not carry the same weight as the Full weekly leaderboard.

YZ Index Smoke快测 AI Evaluation
1,018 08-04

Claude Opus 4.7 Leads at 95.19 Points: 2026-08-03 Smoke Quick Test Data Briefing

On 2026-08-03, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first at 95.19 points. Smoke is a daily 10-question quick test for observing short-term signals, not equivalent to the Full weekly ranking conclusions.

YZ Index Smoke快测 AI Evaluation
1,489 08-03

GLM-4.6 Smoke Evaluation: Main Leaderboard Score 74, Code Execution 82.3, Material Constraint 95, API Failure Leaves Dimensions Missing

GLM-4.6 scored 74.00 on the main leaderboard in today's Smoke evaluation, with 82.30 on code execution and 95.00 on material constraint. Two dimensions are missing due to API failure/timeout and have entered automatic retesting, excluded from this period's ranking.

GLM-4.6 Code Execution Smoke Test
669 08-02

Qwen3 Max Rallies +36.8 to Lead; Gemini 3.1 Pro Slips 5.6 as Biggest Loser

In the July 28–August 2, 2026 Smoke evaluation, Qwen3 Max posted the largest seven-day gain (+36.8) to close at 96.1, while Gemini 3.1 Pro fell 5.6 points to 94.45, making it the biggest loser. The report also analyzes scoring trajectories, volatility drivers, integrity-rating changes, and implications for users.

Qwen3 Max Gemini 3.1 Pro 模型趋势分析
527 08-02

Doubao Pro Leads with 96.7 Points: 2026-08-02 Smoke Quick Test Data Briefing

On August 2, 2026, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first at 96.7 points. The briefing covers daily scores across code execution and material constraint dimensions, along with key fluctuations and integrity signals to monitor.

YZ Index Smoke快测 AI Evaluation
658 08-02

MLPerf Endpoints v0.7: A Foundation Release

MLPerf Endpoints v0.7 marks the foundational release of a buyer-centric AI inference benchmark, publishing initial results from Coreweave, Google, Intel, KRAI, and Nvidia across three benchmarks. The release supports automated submission pipelines, continuous review tooling, and dynamic result visualization, with v1.0 planned for later this year.

MLC MLPerf AI基准测试
925 08-02

SGLang and Miles Add Day-0 Support for Kimi K3

SGLang and Miles Add Day-0 Support for Kimi K3SGLang TeamJuly 27, 2026We are excited to announce Day-0 support for Kimi K3 in SGLang and Miles. K3 is the first open-source model in the 3-trillion-para

LMSYS SGLang Kimi K3
1,032 08-02

Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles

Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in MilesZiang Li, humans& and Miles TeamJuly 29, 2026 TL;DR: We implemented two Blackwell-native RL recipes in Miles: end

LMSYS Blackwell MXFP8
653 08-02

RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUs

RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUsRadixArk & GoogleJuly 30, 2026RadixArk and Google Cloud are partnering to bring SGLang to TPUs, giving developers ultimate fl

LMSYS SGLang TPU
754 08-02

Toward a Cleaner Quantization Stack in SGLang

Toward a Cleaner Quantization Stack in SGLangSGLang X Ascend TeamJuly 28, 2026Quantization has moved from an advanced feature to an essential part of high-throughput LLM serving. As the number of chec

LMSYS SGLang 量化优化
559 08-02

GLM-4.6 Material Constraint Score Plummets 27.3 Points, Main Score Rises 30.2 Points

In today's Smoke evaluation, GLM-4.6's material constraint score dropped from 75.00 to 47.70 points, while its main score rose from 46.29 to 76.47 points.

GLM-4.6 Material Constraints Smoke Test
586 08-01

GPT-o3 Drops 13.9 Points on Today's Main Leaderboard, Losing Ground in Both Code Execution and Material Constraints

GPT-o3 scored 79.28 points on today's Smoke evaluation main leaderboard, down 13.9 points from yesterday's 93.16, with notable declines in both code execution and material constraint dimensions.

GPT-o3 Code Execution Smoke Test
563 08-01

Claude Opus 4.7 and Qwen3 Max Tie at 93.39: 2026-08-01 Smoke Quick Test Data Brief

On 2026-08-01, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and Qwen3 Max tying for first place at 93.39 points. Key signals include GLM-4.6's integrity dropping to warn and multiple models posting sharp overall declines.

YZ Index Smoke快测 AI Evaluation
539 08-01

Qwen3 Max Material Constraint Drops 20 Points to 47.70, Code Execution Surges 37.8 Points, Main Leaderboard Rises 11.8 Points

In today's Smoke evaluation, Qwen3 Max's material constraint score dropped 20 points to 47.70, while code execution soared 37.8 points to 92.50, lifting the main leaderboard by 11.8 points to 72.34.

Qwen3 Max Material Constraints Smoke Test
543 07-31

Grok 4 Code Execution Plunges 19.5 Points, Material Constraint Rises 23.2 Points, Main Leaderboard Drops Only 0.3

In today's Smoke evaluation, Grok 4's code execution score dropped from 92.00 to 72.50, while material constraint rose from 60.90 to 84.10, and the main leaderboard score slightly fell from 78.01 to 77.72.

Grok 4 Code Execution Smoke Test
509 07-31

DeepSeek V4 Pro Leads with 96.94: 2026-07-31 Smoke Quick Test Data Brief

On 2026-07-31, the YZ Index Smoke quick test covered 10 models, with DeepSeek V4 Pro scoring 96.94 to top the daily rankings. The test focuses on code execution and material constraints, serving as a short-term signal indicator.

YZ Index Smoke快测 AI Evaluation
703 07-31
7 8 9 10 11

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0