Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(513) Artificial Intelligence(463) Anthropic(393) AI Safety(299) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(110) xAI(105) Smoke Test(103) Data Centers(102) Code Execution(99) Funding(94) Claude(93) AI(92) Cybersecurity(91) AI Chips(90) Material Constraints(90)

Resource Constraints Become the Hardest Scenario in WDCD, Doubao Scores 3.5 Points in Business Rules, Surpassing GPT

The WDCD five-scenario evaluation reveals that resource constraints is the hardest scenario with the lowest overall scores, while DoubaoPro achieves the highest score in business rules, demonstrating significant model specialization.

WDCD Compliance Test 模型横评
629 05-17

R3 Collapse Rate 93.3%! Grok4 WDCD Three-Round Test: First Round Fully Compliant, Last Round Crashes

The WDCD three-round test reveals that model integrity drops to 30.6% under direct pressure in R3, with Grok4 hitting a 93.3% collapse rate, exposing the fragility of safety alignment.

WDCD Compliance Test 模型衰减
534 05-17

WDCD Commitment Ranking: GPT-5.5 Dominates with 71.67 Points, Grok 4 Trails at 52.5 Points

The WDCD Commitment Test reveals models' true performance under constraints through three rounds of dialogue. GPT-5.5 leads with 71.67 points, while Grok 4 scores only 52.5 points, ranking last—a gap of 19.17 points between the top and bottom.

WDCD Compliance Test AI模型排行
547 05-17

Claude Sonnet 4.6 dropped 12.3 points on main leaderboard, material constraint plummeted 27.3 points in a single day

Claude Sonnet 4.6 showed abnormal results in today's Smoke test, with the material constraint dimension dropping sharply. The drop may be due to sampling variance but warrants further monitoring.

Claude Sonnet 4.6 Material Constraints Smoke Test
502 05-17

Claude Opus 4.7 Smoke Evaluation Main Score Plunges 9 Points, Material Constraint Halves 20 Points in a Single Day

In today's Smoke evaluation, Claude Opus 4.7's main score dropped by 9 points from 97.75 to 88.75, primarily due to a sharp decline in the material constraint dimension from 95 to 75 points—a direct loss of 20 points in a single day.

Claude Opus 4.7 Material Constraints Smoke快测
588 05-17

7-Day Smoke Quick Test: Wenxin Yiyan Soars 53 Points, GPT-o3 Leads with -7.8 Decline

This week's 7-day Smoke Quick Test data reveals polarization: Wenxin Yiyan surged 53.4 points while GPT-o3 fell 7.8 points.

ERNIE Bot GPT-o3 Smoke Test
715 05-17

Three Models Tie at 88.75 for First Place; Claude's Duo Plunges 12 Points; Smoke Rankings Undergo Major Shakeup

Today's Smoke Lite evaluation results show a three-way tie for first place at 88.75 points, while the Claude series suffered sharp declines. The shakeup signals that open models are rapidly closing the gap with closed-source leaders.

Claude Opus 4.7 Material Constraints Smoke Light Test
613 05-17

NTE Game Developer Confirms Ban on AI Core Assets, Community Divided Over Quality vs Efficiency

NTE game development team confirmed that future core assets and character art will not use AI technology, prioritizing quality and reputation. The community is divided over this decision.

AI游戏开发 资产争议 质量优先
967 05-16

Nvidia Releases 2.6B Open-Source World Model: Innovative Breakthrough Sparks Security Controversy

Nvidia has officially released a 2.6B-parameter open-source world model that supports controllable world generation from a single image, text, and trajectory, running on a single GPU. The release has drawn both praise for democratizing AI research and criticism over potential misuse for generating fake content.

NVIDIA 世界模型 AI开源
544 05-16

Anthropic Calls for Aggressive US AI Policy Toward China, Sparks Heated Debate Over Safety Lab Positioning

Anthropic published a new paper on May 14 urging the US government to take more aggressive measures against China in AI. The company's shift from a cautious safety lab to a hawkish stance has sparked intense controversy.

Anthropic AI政策 中美科技
365 05-16

GPT-5.5's Main Ranking Plunges 28 Points: Is It Real Degradation?

GPT-5.5's code execution score dropped from 100 to 50, causing a 28-point drop in the main ranking. But is this degradation or just sampling noise?

GPT-5.5 Code Execution Smoke Test
648 05-16

Gemini 2.5 Pro Drops 10 Points: Ability Intact, Credibility Fails

Gemini 2.5 Pro's credibility rating fell from pass to fail, causing a 10-point drop in the main ranking, even though its code execution score remained perfect.

Gemini 2.5 Pro Material Constraints Smoke Test
623 05-16

Three Models Plunge by 28 Points, Claude Still Near Perfect Score

Today's YZ Index Smoke lightweight test reveals that three leading models suffered significant drops, while Claude models dominate near-perfect scores with structural advantages in code execution and material constraint.

Claude Sonnet 4.6 GPT-5.5 Code Execution
736 05-16

Amazon Launches Shopping-Focused Alexa, E-commerce AI Moves to the Frontline

On May 13, 2026, Amazon launched "Alexa for Shopping," an AI-powered shopping assistant that integrates personalized recommendations, voice purchasing, price comparisons, and deal alerts within the Amazon ecosystem. The move signals a shift in e-commerce from search-based interfaces to conversational AI agents.

Amazon AI购物助手 语音电商
817 05-15

Claude Paid Plans to Include Monthly Usage Credits

Anthropic announced that starting June 15, 2026, Claude paid plans will include monthly credits for programmatic tools like Claude Agent SDK and Claude Code GitHub Actions. This move aims to integrate Claude deeper into development workflows and automation, lowering the barrier for developers to test real-world scenarios.

Claude Anthropic AI开发者工具
3,374 05-15

Meta Launches Meta AI Incognito Chat Mode: Privacy Protection or Data Trade-off?

Meta announced on May 13, 2026, the launch of an incognito chat mode for Meta AI, integrated into WhatsApp and Meta AI apps, allowing private interactions with no data retention. This article analyzes the move from a technical perspective, highlighting strategic shifts and assessing it through the YZ Index v6 methodology.

Meta AI Privacy Protection AI聊天趋势
517 05-15

DeepSeek gains 5 points but fails: 10-question Smoke test alarm

Today's Smoke evaluation shows the main benchmark up by 5 points, but the integrity rating drops from pass to fail, signaling a classic alarm of "seemingly stronger capability but lost trustworthiness at the admission gate."

DeepSeek V4 Pro Integrity Rating Smoke Test
689 05-15

Claude Sonnet 4.6 Material Grounding Plunges 27.5 Points, But Main Leaderboard Rises Against the Trend by 1.4 Points?

In today's Smoke evaluation, Anthropic's Claude Sonnet 4.6 saw a dramatic split: material grounding scores dropped 27.5 points to 69, while code execution surged 25 points to a perfect 100, with the main leaderboard edging up 1.4 points to 86.05.

Claude Sonnet 4.6 Material Constraints Smoke Test
654 05-15

Two Zero-Execution Shocks, Claude Holds at 88.75

Today’s Smoke benchmark shows Claude Opus 4.7 leading with 88.75, while two models scored zero in code execution; the real differentiator is material constraint, not execution ability.

Claude Opus 4.7 Material Constraints Smoke Test
647 05-15

Canada NDP Calls for Moratorium on New AI Data Centers, Sparking Innovation vs. Regulation Conflict

This article evaluates the NDP's proposal for a moratorium on new AI data centers as a policy "product," analyzing its innovations, shortcomings, comparisons, and practical advice. The YZ Index v6 methodology is applied to provide a quantitative assessment.

AI Data Centers 加拿大政策 监管辩论
471 05-14
37 38 39 40 41

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0