Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(513) Artificial Intelligence(461) Anthropic(392) AI Safety(297) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(110) xAI(105) Smoke Test(103) Data Centers(101) Code Execution(99) Funding(94) Claude(93) AI(92) AI Chips(90) Cybersecurity(90) Material Constraints(90)

Claude Sonnet 4.6 Leads with 97.53 Points, Material Constraints Drag ERNIE Bot 40 Points Behind

Smoke's quick test today directly concludes that code execution has become the passing line, while material constraints are the true dividing line. Claude Sonnet 4.6 tops the leaderboard with 97.53 points, followed by Opus 4.7 and Grok 4.

Claude Sonnet 4.6 Material Constraints Smoke Light Test
491 06-10

Apple WWDC 2026: Gemini-Powered Siri Debuts, On-Device AI Reshapes Intelligent Ecosystem

At WWDC 2026, Apple announced Gemini-powered Siri and a multi-model Apple Intelligence architecture, marking a major breakthrough in generative AI.

WWDC Siri Apple Intelligence
400 06-10

OpenAI Secretly Files IPO, AI Giant's Listing Wave Sparks Market Controversy

OpenAI has quietly submitted an IPO filing to the SEC, signaling accelerated commercialization, while its affiliated company Worldcoin reportedly conducts layoffs. This dual development stirs debate in tech and capital markets over the AI industry's transition from innovation to profit-driven expansion.

OpenAI IPO Sam Altman
346 06-10

NVIDIA and Hyundai Deepen AI Collaboration, Accelerating Commercialization of Embodied Intelligent Robots

NVIDIA CEO Jensen Huang recently met with Hyundai Motor Group executives to deepen cooperation in AI applications across mobility, advanced manufacturing, and robotics, marking a new phase in the partnership between global tech giants and traditional automakers in embodied intelligence.

NVIDIA Hyundai Robotics
361 06-10

Moonshot AI Launches $2 Billion Funding Round, Valuation Eyes $30 Billion

Chinese AI startup Moonshot AI has announced a new funding round targeting $2 billion, which would boost its valuation to $30 billion. This marks a major milestone in China's AI sector and reflects sustained investor confidence in generative AI.

Moonshot AI Kimi AI融资
968 06-10

Anthropic Launches Claude Fable 5, Performance Greatly Improved Based on Mythos Architecture

Anthropic recently unveiled the new Claude Fable 5 model, built on the Mythos underlying architecture, marking another major breakthrough in large language models. The model excels in multiple benchmark tests and has attracted widespread developer attention with its affordable pricing.

AI Models Anthropic 产品发布
525 06-10

AI chip stocks plunge $1.3 trillion: Employment data triggers rate hike fears, Nvidia leads decline as market divergence intensifies

AI chip stocks suffered a massive sell-off on Thursday, wiping out approximately $1.3 trillion in market cap. Stronger-than-expected employment data fueled rate hike concerns, with Broadcom's outlook amplifying selling pressure and Nvidia leading the decline, deepening market uncertainty.

AI股票 泡沫争议 Nvidia
435 06-09

OpenAI’s Future Strategy Revealed: Sam Altman Reaffirms AGI for the Benefit of Humanity, Market Discusses Possibility of Government Stake

OpenAI CEO Sam Altman recently unveiled the company’s next-phase strategic plan, with the core goal of ensuring advanced AI technology serves the well-being of all humanity. This statement comes amid multiple lawsuits and debates over the company’s technical direction.

OpenAI AGI 科技战略
632 06-09

Nvidia AI Infrastructure Global Deployment Accelerates: Korean Giants Sign AI Factory Deals, Deepen Robot Collaboration

Nvidia has announced multiple AI infrastructure cooperation agreements with major Korean tech companies, marking further expansion in global AI infrastructure. These partnerships cover AI factory projects, robotics collaboration, and memory supply deals, while Jensen Huang emphasized that current AI-related stock valuations are "very cheap."

Nvidia AI工厂 Robotics
356 06-09

Apple WWDC 2026 Kicks Off: Siri Fully Embraces Gemini Model, AI Deeply Reshapes iOS Ecosystem

At WWDC 2026, Apple announced a comprehensive overhaul of Siri with deep integration of Google's Gemini model, transforming it into a generative AI assistant. The event also introduced AI-powered features in Photos and Shortcuts, signaling a major shift in Apple's AI strategy.

Apple WWDC Siri Gemini AI编辑
6,055 06-09

Smoke Daily: GPT-5.5 tops with 92.58 points, material constraint gap of 19 points decides the outcome

Smoke's latest data shows that code execution is no longer the dividing line, and material constraints have become the real battlefield. A gap of 19.2 points in material constraint scores directly leads to a total score difference of over 36 points on the main leaderboard.

GPT-5.5 Material Constraints 代码执行满分
539 06-09

11 Models Answer Same Blame-Shifting Problem: 8 Get A>B>D>C, 3 Get 0 Points Directly

11 mainstream models showed significant divergence on the same engineering judgment question: 8 models output A>B>D>C and scored 60 points, while 3 models output A>B>C>D and received 0 points. The difference lies only in the relative order of D and C.

execution grounding 工程判断
432 06-08

Binary Tree Serialization Test: 11 Models, 7 Full Scores, 4 Directly Zero

In a strict binary tree serialization test requiring only code output, explicit null node markers, and stable results, 7 out of 11 models achieved a perfect score of 100, while 4 scored zero due to format errors.

Code Execution Material Constraints 二叉树序列化
557 06-08

11 Models Tested on Bracket Matching: 7 Full Scores, 4 Zero Scores

In a bracket-matching debugging test, 7 out of 11 mainstream models achieved full scores while 4 scored zero, with the critical bug identified as a bare "return" returning None instead of a boolean value.

Code Execution Material Constraints 括号匹配
565 06-08

11 AI Models Solve SQL Duplicate Payment Problem: Only 4 Score Full Marks, 7 Score Zero

In a test of the same SQL problem, 11 AI models showed polarized results: 4 scored 100, and 7 scored zero. The core differences lie in self-join deduplication logic, time difference calculation function selection, and the placement of the status condition.

Code Execution Doubao Pro SQL自连接
515 06-08

11 Models All Output [2,2,2] for the Same Closure Problem, Yet All Scored 0 on YZ Index

Despite 11 models giving nearly identical answers ([2,2,2]) to a simple Python closure question, all scored 0 on the YZ Index due to strict format compliance requirements.

Code Execution Material Constraints Python 闭包
513 06-08

GPT-o3 Reservoir Sampling Score Plummets from 100 to 0, Code Execution Truth Hides in Details

In the v6 evaluation, GPT-o3's main score rose from 75.86 to 82.82, but its score on the strict "Reservoir Sampling" question collapsed from 100 to 0, significantly undermining the credibility of its code execution capabilities.

GPT-o3 Code Execution 蓄水池采样
443 06-08

Claude Sonnet 4.6 Drops from 100 to 0 on Strict SQL Question, Yet Main Leaderboard Rises by 9.3

In the v6 evaluation, Claude Sonnet 4.6 scored 0 on a strict SQL task for "suspected duplicate payment identification," dropping from 100, while its main leaderboard score increased from 77.98 to 87.24. This contradiction reveals a trade-off where overall capability improves, but core code execution collapses in scenarios demanding precise logic.

Claude Sonnet 4.6 Code Execution SQL故障
459 06-08

11 Models in Transition: Grok 4 Tops the Charts, DeepSeek Series Exits En Masse

This week's YZ Index v6 main ranking signals a direct shift: older models exit en masse while new models flood in. Among the seven debut models, Qwen3 Max, Grok 4, and ERNIE Bot 4.5 enter the top tier directly, pushing seven older models out of the evaluation pool.

Grok 4 Code Execution 新模型首秀
524 06-08
Research Lab

3 Major Models Translation Showdown: Week 24 Quality Evaluation, passthrough Leads with a Score of 9

This week, <strong>2425</strong> translation tasks were completed by <strong>3</strong> models. <strong>3</strong> samples were selected for multi-model blind comparison, with the overall best being <strong>passthrough</strong> (average score 9/10).

Translation Quality AI Model Comparison passthrough
453 06-08
27 28 29 30 31

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0