Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(670) Artificial Intelligence(559) Anthropic(486) AI Safety(458) AI Agents(225) AI Regulation(174) Meta(170) WDCD(163) Smoke Test(150) Cybersecurity(146) AI Ethics(146) Google(145) Generative AI(131) Code Execution(131) Data Centers(129) Material Constraints(122) Funding(117) Claude(116) xAI(113) AI Chips(111) Compliance Test(109)

US Government's New AI Procurement Rules Rock Silicon Valley: Big Tech Cheers, 95% of Startups May Be Kicked Out

The White House's new "Responsible AI Procurement Executive Order" mandates federal agencies to follow NIST's AI Risk Management Framework when purchasing AI systems, potentially reshaping the industry through market forces while raising concerns about market concentration.

AI Governance 政府采购 合规成本
715 03-24

AlphaFold 3 Emerges: AI Predicts Dynamic Protein Interactions with 90% Accuracy, Pharma Giants' Stock Prices Surge 15%

Google DeepMind's AlphaFold 3 achieves a revolutionary breakthrough in predicting dynamic protein-drug interactions with over 90% accuracy, triggering a massive rally in global biotech stocks while raising questions about the practical challenges of translating computational predictions into clinical applications.

AlphaFold DeepMind 蛋白质折叠
728 03-24

Grok 3 Stability Plummets 22.5 Points: When AI Meets Real Engineering Scenarios, The Truth Comes Out

Grok 3's stability score crashed from 54.2 to 31.7 points in the latest Winzheng evaluation, exposing a fatal weakness in current AI models that excel at coding but fail at real-world engineering judgment.

Grok 3 稳定性测试 工程判断力
1,698 03-22

GPT-o3 Crashes: The Fatal Flaws Behind a 31-Point Plunge

GPT-o3's availability score plummeted from 100 to 69 in just one week, exposing fundamental architectural defects rather than isolated issues—a technical accident that reveals systemic imbalances in AI development.

GPT-o3 可用性测试 模型稳定性
1,164 03-22

GPT-o3 Collapsed: Not Performance Fluctuation, But Systematic Architectural Breakdown

GPT-o3 suffered a catastrophic system failure with stability plummeting from 53 to 28 points and availability dropping from 100 to 69, revealing fundamental architectural flaws rather than typical performance variations.

GPT-o3 稳定性测试 模型架构
975 03-22

GPT-o3 Crashes: 5 Rate Limits in 30 Seconds, Long Context Score Plummets by 33.5 Points

GPT-o3's long context processing capability collapsed in recent testing, with scores dropping from 62.3 to 28.8 points due to aggressive API rate limiting, exposing serious infrastructure issues at OpenAI.

GPT-o3 长上下文 API限流
1,087 03-22

GPT-4o Crashes: The Strict Mode Trap Behind a 35-Point Plunge

GPT-4o experiences a catastrophic performance collapse with its usability score plummeting from 100 to 65, caused by overly conservative "strict tool calling" that makes the model refuse to perform basic tasks.

GPT-4o 可用性测试 严格模式
913 03-22

Technical Risks Behind Doubao Pro's Sharp Decline in Stability

Doubao Pro's stability score plummeted from 54.5 to 34.7 (a 36.3% drop) this week, despite significant improvements in programming and knowledge work dimensions, revealing a concerning pattern of "progress and regression coexisting" that warrants in-depth analysis.

豆包Pro 稳定性测试 AI Evaluation
2,147 03-22

GPT-4o Crashes: 5 Failed Tests Expose OpenAI's Infrastructure Crisis

GPT-4o's catastrophic failure in long-context tests, with 5 questions returning rate limit errors, reveals OpenAI's severe infrastructure problems rather than model capability issues.

GPT-4o 长上下文 OpenAI基础设施
1,000 03-22

Gemini 2.5 Pro Crashes: Engineering Judgment Failure Behind 23-Point Stability Plunge

Gemini 2.5 Pro's stability score plummeted 22.8 points in one week, exposing a critical lack of engineering judgment despite gains in programming capabilities.

Gemini 2.5 Pro 模型稳定性 Google AI
1,510 03-22

Wenxin 4.0 Stability Plummets 22 Points: Why Does Baidu AI Always Drop the Ball at Critical Moments

Wenxin 4.0's stability score crashed from 52.1 to 30 points while programming ability soared by 41.4 points, exposing Baidu's critical engineering shortcomings and raising serious concerns about China's AI industrialization approach.

文心一言4.0 稳定性测试 百度AI
2,640 03-22

Qwen Max Stability Plummets by 22.8 Points: Model Update Triggers Output Quality Volatility

Qwen Max exhibits extreme duality in this week's evaluation, with significant improvements in programming and long-context tasks, but a catastrophic decline in stability metrics. This "fire and ice" performance warrants in-depth analysis.

Qwen Max 稳定性测试 AI Evaluation
1,057 03-22

Technical Concerns Behind Gemini 2.5 Pro's Dramatic Stability Decline

This week's evaluation data reveals Gemini 2.5 Pro's stability score plummeted from 54.0 to 31.2, a 42.2% drop, exposing serious issues in maintaining consistent output quality while other metrics improved.

Gemini 模型稳定性 性能评测
1,814 03-22

DeepSeek R1 Stability Plummets 22 Points: The Truth Behind Complete Failure on Simple Judgment Questions

DeepSeek R1's stability score crashed from 53.7 to 31.6 points this week, with the model failing basic judgment questions like whether water can boil at 101°C under standard pressure, raising serious concerns about its reliability.

DeepSeek R1 稳定性测试 AI推理失败
916 03-22

Claude 4.6 Version Crashes: The Algorithmic Black Hole Behind a 23-Point Plunge

While everyone celebrates Claude's 38.3-point programming improvement, a more dangerous signal has been masked: stability plummeted from 54.2 to 31.2 points, revealing a systemic algorithmic collapse rather than normal performance fluctuation.

Claude 稳定性测试 Model Degradation
1,144 03-22

Technical Risks Behind Wenxin Yiyan 4.0's 22-Point Stability Plunge

Wenxin Yiyan 4.0 showed remarkable anomalies in this week's evaluation, with programming capability surging 41.4 points but stability plummeting from 52.1 to 30.0 points, revealing potential deep-seated issues in the model upgrade process.

ERNIE Bot 模型稳定性 性能评测
793 03-22

Technical Analysis: DeepSeek V3's Stability Plunges by 21.4 Points

DeepSeek V3 shows contradictory performance this week with programming capabilities soaring 42.6 points while stability metrics collapse from 53.4 to 32.0 points, revealing critical trade-offs in AI model optimization.

DeepSeek V3 稳定性测试 Model Evaluation
803 03-22

11 AI Models Surge 40 Points in Programming Tests: What Really Happened?

A massive collective surge in AI model programming scores reveals hidden signals about the industry, including Chinese models dominating rankings for the first time and OpenAI's concerning decline.

DeepSeek GPT-o3 编程能力测试
944 03-22

Technical Concerns Behind DeepSeek R1's 22-Point Stability Plunge

DeepSeek R1 shows extreme performance polarization in this week's evaluation: programming capability soared 47.4 points while stability plummeted 22.1 points, revealing critical trade-offs in model optimization.

DeepSeek R1 稳定性测试 Model Evaluation
926 03-22

The Technical Truth Behind Claude 3.5 Sonnet's 23-Point Stability Plunge

Claude 3.5 Sonnet (version 4.6) experienced a dramatic 42% drop in stability scores from 54.2 to 31.2, while simultaneously achieving significant improvements in programming capabilities and other dimensions, suggesting aggressive optimization strategies that may have compromised output consistency.

Claude 稳定性测试 AI Benchmarks
1,063 03-22
71 72 73 74 75

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷 XunOPC

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0