Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(697) Artificial Intelligence(566) Anthropic(512) AI Safety(510) AI Agents(234) AI Regulation(200) Meta(176) WDCD(173) Smoke Test(157) Cybersecurity(152) Google(151) AI Ethics(149) Generative AI(140) Data Centers(138) Code Execution(134) Material Constraints(126) Funding(124) Claude(119) Compliance Test(117) AI Chips(115) xAI(114)

Higgs Audio v3 TTS on SGLang-Omni: Real-Time, Controllable Speech for Voice Agents

Higgs Audio v3 TTS on SGLang-Omni: Real-Time, Controllable Speech for Voice AgentsBoson AI & SGLang-Omni TeamJune 4, 2026Today we are announcing end-to-end serving for Higgs Audio v3 TTS on…

LMSYS TTS语音合成 多语言模型
441 06-29

Announcing the Recipient of the 2026 LMSYS PhD Fellowship

Announcing the Recipient of the 2026 LMSYS PhD FellowshipLMSYS BoardJune 08, 2026We are delighted to announce the first recipient of the LMSYS Fellowship Program: Will Lin. Following the launch of…

LMSYS 博士奖学金 Open Source AI
539 06-29

The next generation of speculative decoding: DFlash and Spec V2

The next generation of speculative decoding: DFlash and Spec V2Z Lab, Modal, and SGLang TeamsJune 15, 2026Using Modal and Z Lab's DFlash speculative decoding models with SGLang’s newly default Spec…

LMSYS 推测解码 DFlash
592 06-29

MOSS-TTS Local Transformer v1.5 on SGLang-Omni: Serving Native-Streaming 48 kHz Speech

MOSS-TTS Local Transformer v1.5 on SGLang-Omni: Serving Native-Streaming 48 kHz SpeechMOSI, OpenMOSS Team & SGLang-Omni TeamJune 17, 2026Today we are announcing end-to-end serving for…

LMSYS TTS模型 语音合成
506 06-29

Optimizing Ling-2.6-1T on TPU with SGLang-JAX: Hiding MoE Data Movement Behind Compute with One Pallas Kernel

Optimizing Ling-2.6-1T on TPU with SGLang-JAX: Hiding MoE Data Movement Behind Compute with One Pallas KernelPrayer, JamesBrianD, Haolin Fu, Haoguang Cai, Qinghan ChenJune 17, 2026SGLang-JAX now…

LMSYS MoE 优化 TPU 推理
412 06-29

Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLB

Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLBNVIDIA TeamJune 26, 2026TL;DR Mixture-of-Experts (MoE) models rely on Expert Parallelism (EP) to scale inference across multiple…

LMSYS MoE SGLang
548 06-29

Doubao Pro Smoke Evaluation Main Ranking Drops 13.8 Points, Code Execution Falls from 100 to 75

In the June 2026 YZ Index evaluation of 11 models, Doubao Pro's main ranking score dropped from 98.61 yesterday to 84.77 today, a decline of 13.8 points. The code execution dimension fell 25 points from 100.00 to 75.00, accounting for almost the entire drop.

Doubao Pro 主榜 Smoke测试
588 06-29

Claude Opus 4.7 Tops Main Leaderboard with Perfect 100 Points, Doubao Pro Plunges 13.8 Points Exposing Execution Weakness

In the June 29, 2026 YZ Index Smoke Lite test, Claude Opus 4.7 ranked first with a perfect 100 on the main leaderboard, 100 on execution, and 100 on compliance [pass], achieving full marks in both execution and compliance.

Claude Opus 4.7 Doubao Pro 执行约束
375 06-29

Claude Scores Largest Increase of 19.8 Points; All Eight WDCD Models Rise, None Decline

In the latest WDCD cycle (Run #196), all eight evaluated models showed positive changes, with none declining. Claude Opus 4.7 recorded the largest single-model increase of 19.8 points, jumping to 89.29 points and entering the top three.

WDCD Compliance Test 模型性能变化
685 06-28

WDCD Review: Safety Compliance Becomes the Biggest Weakness, Highest Score Among 11 Models Only 3.57

In the WDCD Compliance Test, the safety compliance scenario scored the lowest on average across all models, with the highest score being only 3.57/4 for deepseek-v4-pro, while claude-sonnet-4.6 scored only 2.57/4.

WDCD Compliance Test 安全合规
641 06-28

Grok 4 Zero Crashes Overwhelms GPT-o3's 17% Collapse: WDCD Three-Round Attenuation Reveals True Resilience

In WDCD tests, Grok 4 maintains a 1.83/2 honesty rate with zero crashes in R3, while both Claude Sonnet 4.6 and GPT-o3 suffer six complete R3 crashes each (17.1%). The data reveals systematic degradation across three pressure rounds, with overall honesty dropping to 1.63/2 in the final stage.

WDCD Compliance Test 三轮衰减
883 06-28

Gemini 3.1 Pro Scores 93.57 Points, Tops WDCD Compliance Rankings; ERNIE Bot 4.5 Only 75.71 Points, Last Place

Gemini 3.1 Pro leads the WDCD compliance ranking with 93.57 points (R1=1.00, R2=0.97, R3=1.77/2), while ERNIE Bot 4.5 ranks 11th with 75.71 points (R1=0.89, R2=0.60, R3=1.54/2).

WDCD Compliance Test 排行榜分析
566 06-28

Claude Sonnet 4.6 Smoke Review Main Score Plummets 25.9 Points, Code Execution Drops from 100 to 50

In the June 2026 Smoke review of the YZ Index, Claude Sonnet 4.6's main score fell from 96.45 to 70.52, code execution dropped from 100.00 to 50.00, while material constraint rose from 92.10 to 95.60.

Claude Sonnet 4.6 Code Execution Smoke Test
632 06-28

Claude Opus 4.7 Code Execution Plummets from 100 to 50, Main Score Drops 25.7 Points in a Single Day

In today's Smoke evaluation of the YZ Index, Claude Opus 4.7 saw its main score drop from 97.12 to 71.47, a decline of 25.7 points, driven entirely by the code execution dimension halving from 100.00 to 50.00.

Claude Opus 4.7 Code Execution Smoke Test
571 06-28

YZ Index Smoke Weekly: ERNIE Bot 4.5 Drops 37.2 Points, Multiple Models Fluctuate Over 28

In the YZ Index Smoke tests from June 23 to 28, 2026, ERNIE Bot 4.5 showed the largest decline, dropping 37.2 points from 98.74 to 61.52, with an average of 82.1, while multiple models experienced fluctuations exceeding 28 points. Only Doubao Pro maintained a slight upward trend, emerging as the sole stable performer.

文心一言 4.5 Claude Sonnet 4.6 Smoke测试
424 06-28

Doubao Pro tops Smoke benchmark with 98.61 points, Claude's Execution plummets to 50 points

In the Smoke lightweight benchmark on June 28, 2026, Doubao Pro topped the main leaderboard with 98.61 points (Execution 100, Material Constraint 96.9). Claude Opus 4.7 and Sonnet 4.6 saw their Execution scores drop from 100 to 50, causing significant ranking declines.

Doubao Pro Claude Opus 执行维度
504 06-28

Claude Opus 4.7 Leads with 97.12 Points, Perfect Execution but Material Constraint Score of 93.6 Drags Down Overall

In the YZ Index from June 27, 2026, Smoke lightweight evaluation, Claude Opus 4.7 ranked first on the main leaderboard with 97.12 points, achieving a perfect 100 points in code execution and 93.6 points in material constraint.

Claude Opus 4.7 Code Execution Smoke Light Test
558 06-27

Qwen3 Max Code Execution Plunges 50 Points, Main Ranking Only Drops 1.5 Points

In the June 2026 YZ Index evaluation of 11 models, Qwen3 Max's code execution score plummeted from 100.00 to 50.00 in a single day. However, the main ranking only dropped by 1.5 points, as gains in material adherence and side metrics offset the decline.

Qwen3 Max Code Execution 烟雾测试
654 06-24

Claude Opus 4.7 Smoke Evaluation Main Benchmark Drops 27.5 Points, Code Execution from 100 to 50

In the June 2026 YZ Index test of 11 models, Claude Opus 4.7 Smoke's main benchmark score dropped from 100.00 yesterday to 72.50 today, with the code execution dimension falling directly from 100.00 to 50.00.

Claude Opus 4.7 Code Execution Smoke快测
701 06-24

4 Models' Execution Scores Plummet to 50, ERNIE Bot's Main Leaderboard Drops 34.1 Points

In the YZ Index Smoke Lightweight Evaluation on June 24, 2026, ERNIE Bot 4.5's main leaderboard score plunged 34.1 points to 64.63, with its execution dimension dropping directly from 100 to 50.

Code Execution Material Constraints ERNIE Bot 4.5
786 06-24
15 16 17 18 19

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC Xinyuan luo

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0