Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(480) Artificial Intelligence(447) Anthropic(372) AI Safety(257) AI Agents(176) Meta(127) AI Ethics(122) WDCD(112) AI Regulation(111) Google(108) Generative AI(107) xAI(102) Data Centers(97) Code Execution(94) Smoke Test(94) Claude(90) AI Chips(90) Funding(87) AI(84) Material Constraints(84) Elon Musk(80)

Behind Google's Reorganization: The Power Play Between Centralization and Decentralization in AI R&D

Alphabet CEO announces the integration of DeepMind, Google Brain and other AI teams into an independent "Google AI" division led by Demis Hassabis, marking the company's largest reorganization that reflects a fundamental shift in AI research and development models.

Google重组 AI战略 DeepMind
632 03-22

NVIDIA B200 GPU In-Depth Review: A Computational Revolution for the AGI Era or Just Overhyped Marketing?

NVIDIA unveils the B200 'Blackwell Ultra' GPU at GTC 2026, featuring 2nm process technology and claiming 30x performance improvement over H100. While the hardware represents a significant leap for AGI-scale models, questions remain about yield rates, real-world performance, and whether the industry truly needs such extreme computational power yet.

NVIDIA B200 GPU AI硬件
1,362 03-22

11 AIs Answer the Same Debugging Question: 5 Score Zero, Where's the Fatal Gap?

Testing 11 mainstream AI models with a real debugging scenario revealed that 45% couldn't even pass, including the newly released DeepSeek V3. The test exposed three critical blind spots in current AI models when handling engineering problems.

豆包Pro Claude 工程调试
1,171 03-21

11 AIs Answer the Same Question, 6 Get Even the Day of the Week Wrong

A simple time zone calculation that elementary school students can solve exposed the shocking reality: over half of top AI models failed completely, and none recognized that March 15th falls during US Daylight Saving Time.

DeepSeek GPT-4o 时区计算
870 03-21

11 AIs Tackle the Same Logic Puzzle, 3 Failures Expose Reasoning Black Holes

A simple logic puzzle involving 5 people's rankings stumped 3 out of 11 AI models, including DeepSeek V3 and Grok 3, revealing fundamental weaknesses in current AI reasoning capabilities despite their acclaimed performance on complex tasks.

DeepSeek Grok 逻辑推理
1,229 03-21

11 AIs Answer Same Question: Doubao Scores 100, 8 Models Score 0

When given the same engineering judgment question, Doubao Pro scored perfect 100 while 8 major AI models including Claude and GPT-4o scored 0, revealing a stark divide in practical problem-solving abilities.

豆包Pro 工程判断力 群发功能调试
922 03-21

When 11 AIs Answer the Same Question, Only 1 Discovers the Truth: The Code Has No Bug

A Python code that ran smoothly for 6 months suddenly threw an error. When 11 top AI models were asked to find the bug, only one discovered the truth: there was no bug in the code at all.

GPT-o3 Claude AI测试
1,060 03-21

11 AIs Answer the Same Question, 10 Are Playing Dumb: Why Did Doubao Get a Perfect Score?

A simple server configuration verification test revealed that 10 out of 11 leading AI models, including GPT-4o and Claude, gave perfunctory responses, while only Doubao Pro provided a comprehensive, practical solution that addressed the real workplace scenario.

豆包 DeepSeek 工程思维
611 03-21

11 AIs Answer the Same Question, 7 Fail: Who's Pretending to Be Smart?

A real-world engineering scenario exposed that over 60% of top AI models prioritize reporting over immediate action during data breaches, with Chinese models surprisingly outperforming their Western counterparts.

DeepSeek Claude 安全事件响应
833 03-21

Grok 3's Logic Score Plummets to Zero: Five Letters Expose Fatal Algorithm Flaw

Grok 3's logic reasoning score collapsed from 100 to 0 in the latest YZ Index evaluation, exposing a systemic failure in the model's reasoning capabilities despite improvements in other areas.

Grok 3 逻辑推理 Model Evaluation
710 03-21

GPT-4o Crashes: Engineers' Most Trusted AI's Judgment Drops to 0

GPT-4o's bug detection capability catastrophically failed in the latest evaluation, scoring 0 on a basic code review test while paradoxically improving its overall programming score, revealing systemic issues in AI development priorities.

GPT-4o 编程能力 代码审查
612 03-21

GPT-4o's Zero-Score Crash on Strict Test: When AI Meets the Friday Deployment Death Trap

GPT-4o catastrophically failed a real-world engineering judgment test about Friday deployments, exposing a critical gap between technical capability and practical engineering wisdom.

GPT-4o 工程判断力 周五发布
636 03-21

Gemini 2.5 Pro's Judgment Hits Zero: Choosing to Report P0 Security Incident Instead of Taking Action

Gemini 2.5 Pro scored 0 on engineering judgment when faced with a critical data breach scenario, exposing a fundamental flaw in AI decision-making during emergencies.

Gemini 2.5 Pro 工程判断力 数据安全事故
721 03-21

Gemini 2.5 Pro Scores 0 from 100 on Time Zone Reasoning: How Terrifying Are LLMs' Common Sense Blind Spots

A simple time zone question that elementary school students can answer correctly caused Google's most powerful model Gemini 2.5 Pro to fail completely, exposing systematic deficiencies in LLMs' handling of real-world basic common sense.

Gemini 2.5 Pro 严格题测试 时区推理
671 03-21

Wenxin 4.0's One Line of Code Exposes Fatal Flaw: When AI Can't Even Recognize a Dictionary

Wenxin 4.0, Baidu's GPT-4 competitor, failed catastrophically on a basic Python dictionary comprehension task, outputting a list instead of a dictionary along with mysterious extra numbers, raising serious concerns about the stability of Chinese AI models.

文心一言4.0 编程能力 代码生成
897 03-21

Doubao Pro Scores Zero on Perfect Question: Why AI Models Collectively Fall Silent During Real Security Incidents

Doubao Pro failed catastrophically on a previously perfect security response question, exposing a fatal flaw in AI decision-making during critical moments. The model prioritized evidence preservation over damage control during a simulated breach, revealing systemic issues in how AI handles real-world emergency scenarios.

豆包Pro 工程判断力 安全事件响应
788 03-21

Claude 4.6 Crashes: The Fatal Flaw Behind Complete Failure on 100-Point Security Questions

Claude Opus 4.6's complete failure on security response questions reveals a deeper issue: when AI encounters real emergencies, their "perfect answers" might be the most dangerous.

Claude Opus 4.6 工程判断力 安全事件响应
747 03-21

Behind GPT-o3's 8.7-Point Surge: Weekly Testing of 11 AI Models Reveals 3 Dangerous Signals

Weekly testing of 11 top AI models reveals concerning trends: stability has become a luxury, long-context capabilities are collectively declining, and Chinese models are reshaping the competitive landscape.

GPT-o3 豆包Pro 模型稳定性
585 03-21

Sora 2.0: The Double-Edged Sword of Generative AI and Regulatory Challenges

Sora 2.0's powerful video generation capabilities have sparked global attention, showcasing both the tremendous potential for creative industries and concerns about misinformation proliferation, presenting a critical test for technology regulation.

Generative AI 虚假信息 技术监管
584 03-21

Meta Llama 4 Open Source Sparks Safety Debate: AI Democratization or Global Risk?

Meta's open-sourcing of Llama 4 on GitHub has ignited fierce debate between developers celebrating AI democratization and security experts warning of weaponization risks. The controversy reveals deeper geopolitical tensions and governance gaps in the AI landscape.

AI开源 Llama4 Meta
999 03-21
47 48 49 50 51

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0