Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(512) Artificial Intelligence(460) Anthropic(391) AI Safety(296) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(110) xAI(105) Smoke Test(103) Data Centers(100) Code Execution(99) Funding(94) Claude(93) AI(92) AI Chips(90) Cybersecurity(90) Material Constraints(90)
Research Lab

WDCD Run #202: Average Instruction Decay Hits -73.2% Across 11 Models, Gemini 3.1 Pro Leads

WDCD Run #202 (2026-06-28) measured multi-turn commitment integrity across 11 frontier models, recording an average instruction decay of -73.2% between Round 1 and Round 3. Gemini 3.1 Pro topped the leaderboard at 93.6 points.

WDCD AI benchmark instruction decay
336 06-28

Claude Scores Largest Increase of 19.8 Points; All Eight WDCD Models Rise, None Decline

In the latest WDCD cycle (Run #196), all eight evaluated models showed positive changes, with none declining. Claude Opus 4.7 recorded the largest single-model increase of 19.8 points, jumping to 89.29 points and entering the top three.

WDCD Compliance Test 模型性能变化
377 06-28

WDCD Review: Safety Compliance Becomes the Biggest Weakness, Highest Score Among 11 Models Only 3.57

In the WDCD Compliance Test, the safety compliance scenario scored the lowest on average across all models, with the highest score being only 3.57/4 for deepseek-v4-pro, while claude-sonnet-4.6 scored only 2.57/4.

WDCD Compliance Test 安全合规
301 06-28

Grok 4 Zero Crashes Overwhelms GPT-o3's 17% Collapse: WDCD Three-Round Attenuation Reveals True Resilience

In WDCD tests, Grok 4 maintains a 1.83/2 honesty rate with zero crashes in R3, while both Claude Sonnet 4.6 and GPT-o3 suffer six complete R3 crashes each (17.1%). The data reveals systematic degradation across three pressure rounds, with overall honesty dropping to 1.63/2 in the final stage.

WDCD Compliance Test 三轮衰减
577 06-28

Gemini 3.1 Pro Scores 93.57 Points, Tops WDCD Compliance Rankings; ERNIE Bot 4.5 Only 75.71 Points, Last Place

Gemini 3.1 Pro leads the WDCD compliance ranking with 93.57 points (R1=1.00, R2=0.97, R3=1.77/2), while ERNIE Bot 4.5 ranks 11th with 75.71 points (R1=0.89, R2=0.60, R3=1.54/2).

WDCD Compliance Test 排行榜分析
281 06-28

Claude Sonnet 4.6 Smoke Review Main Score Plummets 25.9 Points, Code Execution Drops from 100 to 50

In the June 2026 Smoke review of the YZ Index, Claude Sonnet 4.6's main score fell from 96.45 to 70.52, code execution dropped from 100.00 to 50.00, while material constraint rose from 92.10 to 95.60.

Claude Sonnet 4.6 Code Execution Smoke Test
314 06-28

Claude Opus 4.7 Code Execution Plummets from 100 to 50, Main Score Drops 25.7 Points in a Single Day

In today's Smoke evaluation of the YZ Index, Claude Opus 4.7 saw its main score drop from 97.12 to 71.47, a decline of 25.7 points, driven entirely by the code execution dimension halving from 100.00 to 50.00.

Claude Opus 4.7 Code Execution Smoke Test
268 06-28

YZ Index Smoke Weekly: ERNIE Bot 4.5 Drops 37.2 Points, Multiple Models Fluctuate Over 28

In the YZ Index Smoke tests from June 23 to 28, 2026, ERNIE Bot 4.5 showed the largest decline, dropping 37.2 points from 98.74 to 61.52, with an average of 82.1, while multiple models experienced fluctuations exceeding 28 points. Only Doubao Pro maintained a slight upward trend, emerging as the sole stable performer.

文心一言 4.5 Claude Sonnet 4.6 Smoke测试
186 06-28

Doubao Pro tops Smoke benchmark with 98.61 points, Claude's Execution plummets to 50 points

In the Smoke lightweight benchmark on June 28, 2026, Doubao Pro topped the main leaderboard with 98.61 points (Execution 100, Material Constraint 96.9). Claude Opus 4.7 and Sonnet 4.6 saw their Execution scores drop from 100 to 50, causing significant ranking declines.

Doubao Pro Claude Opus 执行维度
268 06-28

OpenAI and Broadcom Unveil Jalapeño Chip: Targeting 50% Inference Cost Reduction, Training Still Relies on NVIDIA

OpenAI and Broadcom jointly announced the launch of Jalapeño, the first custom ASIC chip optimized for large language model inference, aiming to reduce per-response costs by approximately 50% and lessen reliance on NVIDIA. Scheduled for deployment by end of 2026 and mass production by 2027-2028.

AI Chips OpenAI Broadcom
1,311 06-27

Anthropic Accuses Alibaba of Extracting Claude Using 25,000 Accounts; No Public Response Yet

Anthropic sent letters to Reuters and the U.S. Congress on June 24-25, 2026, accusing Alibaba affiliates of using approximately 25,000 fake accounts to generate over 28.8 million Claude interactions from April 22 to June 5, attempting to distill its reasoning and programming capabilities.

AI模型安全 数据提取争议 中美AI竞争
438 06-27

Hasbro Requires Peppa Pig Child Stars to Sign AI Voice Licensing Clause; British Child Star Agency Association Publicly Opposes

Hasbro has added an AI voice replication clause to Peppa Pig renewal contracts, requiring child stars to agree to permanent use of their voices for commercial assets. The British Child Star Agency Association has publicly opposed this.

AI配音权 童星合同 娱乐行业
278 06-27

AI Data Center Demand Explodes: Micron Earnings Lead Chip Stocks Higher

The artificial intelligence wave is profoundly reshaping the supply-demand landscape of the semiconductor industry. Following strong earnings from memory chip giant Micron Technology, related chip stocks have surged, reigniting market enthusiasm for AI data center construction.

芯片股 AI需求 Micron
220 06-27

AI Agent Tools Explode: Claude Sales and Others Lead a New Trend in Sales Automation

The rise of AI agents is transforming sales automation, with tools like Claude Sales enabling one-click outreach and automated follow-ups. The trend marks a shift from model training to practical application, though concerns over data privacy and reliability remain.

AI Agents Claude 销售自动化
240 06-27

Anthropic Employees' Inner Voice: AI Automation Triggers Self-Doubt, What's the Meaning of Human Work?

Several employees of AI company Anthropic have publicly shared psychological distress after using the company's AI product Claude, questioning the value of their own work and experiencing mild depressive symptoms. This phenomenon has sparked heated debate in the tech industry about AI replacing the meaning of human work.

Anthropic AI取代 员工心态
214 06-27

OpenAI Delays Public Release of GPT-5.6, U.S. Government Security Review Sparks AI Regulation Controversy

OpenAI recently announced delaying the public release of its GPT-5.6 model due to U.S. government national security concerns over frontier AI technology, igniting widespread debate on AI regulation and innovation freedom.

OpenAI GPT-5.6 AI Regulation
274 06-27

OpenAI and Broadcom Launch First Custom AI Inference Chip, Expected to Cut Costs by 50%

OpenAI has officially announced the launch of its first custom AI inference chip in partnership with semiconductor giant Broadcom, a breakthrough expected to reduce data center operational costs by up to 50% and support larger-scale AI model deployments.

OpenAI Broadcom AI Chips
218 06-27

Claude Opus 4.7 Leads with 97.12 Points, Perfect Execution but Material Constraint Score of 93.6 Drags Down Overall

In the YZ Index from June 27, 2026, Smoke lightweight evaluation, Claude Opus 4.7 ranked first on the main leaderboard with 97.12 points, achieving a perfect 100 points in code execution and 93.6 points in material constraint.

Claude Opus 4.7 Code Execution Smoke Light Test
272 06-27

OpenAI GPT-5.6 Preview Released in Batches Due to Government Review, Safety vs Innovation Debate Intensifies

On June 25, 2026, OpenAI confirmed that GPT-5.6 would only be available to a small group of partners, each subject to individual U.S. government review, directly limiting early access to the model and fueling debate between safety and innovation advocates.

OpenAI GPT-5.6 AI Regulation
2,005 06-26

U.S. Government Suspends Anthropic Fable 5 Model: Security Review and AI Competitiveness Conflict Intensifies

From June 24 to 25, 2026, the U.S. government demanded that Anthropic fully suspend global access to Fable 5 and Mythos models on grounds of export controls and national security, escalating tensions between safety oversight and AI competitiveness.

AI Regulation Anthropic Fable 5
395 06-26
17 18 19 20 21

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0