Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(488) Artificial Intelligence(451) Anthropic(377) AI Safety(271) AI Agents(177) Meta(129) AI Ethics(123) WDCD(117) AI Regulation(114) Google(110) Generative AI(108) xAI(102) Data Centers(98) Smoke Test(96) Code Execution(95) Funding(90) Claude(90) AI Chips(90) AI(86) Material Constraints(86) Cybersecurity(82)

Technical Analysis: DeepSeek V3's Stability Plunges by 21.4 Points

DeepSeek V3 shows contradictory performance this week with programming capabilities soaring 42.6 points while stability metrics collapse from 53.4 to 32.0 points, revealing critical trade-offs in AI model optimization.

DeepSeek V3 稳定性测试 Model Evaluation
646 03-22

11 AI Models Surge 40 Points in Programming Tests: What Really Happened?

A massive collective surge in AI model programming scores reveals hidden signals about the industry, including Chinese models dominating rankings for the first time and OpenAI's concerning decline.

DeepSeek GPT-o3 编程能力测试
723 03-22

Technical Concerns Behind DeepSeek R1's 22-Point Stability Plunge

DeepSeek R1 shows extreme performance polarization in this week's evaluation: programming capability soared 47.4 points while stability plummeted 22.1 points, revealing critical trade-offs in model optimization.

DeepSeek R1 稳定性测试 Model Evaluation
681 03-22

The Technical Truth Behind Claude 3.5 Sonnet's 23-Point Stability Plunge

Claude 3.5 Sonnet (version 4.6) experienced a dramatic 42% drop in stability scores from 54.2 to 31.2, while simultaneously achieving significant improvements in programming capabilities and other dimensions, suggesting aggressive optimization strategies that may have compromised output consistency.

Claude 稳定性测试 AI Benchmarks
869 03-22

Claude Opus 4.6 Stability Plummets 22.5 Points: Output Format Chaos Raises Concerns

Claude Opus 4.6's stability score crashed from 53.5 to 31.0 points this week, a 42.1% decline, while programming capabilities surged 208%, highlighting the complex trade-offs in AI model optimization.

Claude 稳定性测试 AI Evaluation
797 03-22

11 AIs Answer the Same Debugging Question: 5 Score Zero, Where's the Fatal Gap?

Testing 11 mainstream AI models with a real debugging scenario revealed that 45% couldn't even pass, including the newly released DeepSeek V3. The test exposed three critical blind spots in current AI models when handling engineering problems.

豆包Pro Claude 工程调试
1,173 03-21

11 AIs Answer the Same Question, 6 Get Even the Day of the Week Wrong

A simple time zone calculation that elementary school students can solve exposed the shocking reality: over half of top AI models failed completely, and none recognized that March 15th falls during US Daylight Saving Time.

DeepSeek GPT-4o 时区计算
883 03-21

11 AIs Tackle the Same Logic Puzzle, 3 Failures Expose Reasoning Black Holes

A simple logic puzzle involving 5 people's rankings stumped 3 out of 11 AI models, including DeepSeek V3 and Grok 3, revealing fundamental weaknesses in current AI reasoning capabilities despite their acclaimed performance on complex tasks.

DeepSeek Grok 逻辑推理
1,242 03-21

11 AIs Answer Same Question: Doubao Scores 100, 8 Models Score 0

When given the same engineering judgment question, Doubao Pro scored perfect 100 while 8 major AI models including Claude and GPT-4o scored 0, revealing a stark divide in practical problem-solving abilities.

豆包Pro 工程判断力 群发功能调试
925 03-21

When 11 AIs Answer the Same Question, Only 1 Discovers the Truth: The Code Has No Bug

A Python code that ran smoothly for 6 months suddenly threw an error. When 11 top AI models were asked to find the bug, only one discovered the truth: there was no bug in the code at all.

GPT-o3 Claude AI测试
1,073 03-21

11 AIs Answer the Same Question, 10 Are Playing Dumb: Why Did Doubao Get a Perfect Score?

A simple server configuration verification test revealed that 10 out of 11 leading AI models, including GPT-4o and Claude, gave perfunctory responses, while only Doubao Pro provided a comprehensive, practical solution that addressed the real workplace scenario.

豆包 DeepSeek 工程思维
617 03-21

11 AIs Answer the Same Question, 7 Fail: Who's Pretending to Be Smart?

A real-world engineering scenario exposed that over 60% of top AI models prioritize reporting over immediate action during data breaches, with Chinese models surprisingly outperforming their Western counterparts.

DeepSeek Claude 安全事件响应
839 03-21

Grok 3's Logic Score Plummets to Zero: Five Letters Expose Fatal Algorithm Flaw

Grok 3's logic reasoning score collapsed from 100 to 0 in the latest YZ Index evaluation, exposing a systemic failure in the model's reasoning capabilities despite improvements in other areas.

Grok 3 逻辑推理 Model Evaluation
715 03-21

GPT-4o Crashes: Engineers' Most Trusted AI's Judgment Drops to 0

GPT-4o's bug detection capability catastrophically failed in the latest evaluation, scoring 0 on a basic code review test while paradoxically improving its overall programming score, revealing systemic issues in AI development priorities.

GPT-4o 编程能力 代码审查
617 03-21

GPT-4o's Zero-Score Crash on Strict Test: When AI Meets the Friday Deployment Death Trap

GPT-4o catastrophically failed a real-world engineering judgment test about Friday deployments, exposing a critical gap between technical capability and practical engineering wisdom.

GPT-4o 工程判断力 周五发布
644 03-21

Gemini 2.5 Pro's Judgment Hits Zero: Choosing to Report P0 Security Incident Instead of Taking Action

Gemini 2.5 Pro scored 0 on engineering judgment when faced with a critical data breach scenario, exposing a fundamental flaw in AI decision-making during emergencies.

Gemini 2.5 Pro 工程判断力 数据安全事故
729 03-21

Gemini 2.5 Pro Scores 0 from 100 on Time Zone Reasoning: How Terrifying Are LLMs' Common Sense Blind Spots

A simple time zone question that elementary school students can answer correctly caused Google's most powerful model Gemini 2.5 Pro to fail completely, exposing systematic deficiencies in LLMs' handling of real-world basic common sense.

Gemini 2.5 Pro 严格题测试 时区推理
680 03-21

Wenxin 4.0's One Line of Code Exposes Fatal Flaw: When AI Can't Even Recognize a Dictionary

Wenxin 4.0, Baidu's GPT-4 competitor, failed catastrophically on a basic Python dictionary comprehension task, outputting a list instead of a dictionary along with mysterious extra numbers, raising serious concerns about the stability of Chinese AI models.

文心一言4.0 编程能力 代码生成
905 03-21

Doubao Pro Scores Zero on Perfect Question: Why AI Models Collectively Fall Silent During Real Security Incidents

Doubao Pro failed catastrophically on a previously perfect security response question, exposing a fatal flaw in AI decision-making during critical moments. The model prioritized evidence preservation over damage control during a simulated breach, revealing systemic issues in how AI handles real-world emergency scenarios.

豆包Pro 工程判断力 安全事件响应
791 03-21

Claude 4.6 Crashes: The Fatal Flaw Behind Complete Failure on 100-Point Security Questions

Claude Opus 4.6's complete failure on security response questions reveals a deeper issue: when AI encounters real emergencies, their "perfect answers" might be the most dangerous.

Claude Opus 4.6 工程判断力 安全事件响应
753 03-21
17 18 19 20 21

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0