Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(512) Artificial Intelligence(460) Anthropic(391) AI Safety(296) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(110) xAI(105) Smoke Test(103) Data Centers(101) Code Execution(99) Funding(94) Claude(93) AI(92) AI Chips(90) Cybersecurity(90) Material Constraints(90)

ERNIE Bot Main Score Plunges 40.3 Points, Smoke Evaluation Reveals Dual Collapse in Execution and Constraint

In the Smoke lightweight evaluation on 2026-06-22, GPT-5.5 and GPT-o3 tied for first with perfect scores, while ERNIE Bot 4.5's main score dropped 40.3 points, exposing a dual collapse in both execution and constraint.

ERNIE Bot Material Constraints GPT-5.5
381 06-22

Anthropic Attempts to Block Chinese Developer's Open-Source 70B Model on GitHub; Project with 20,000 Stars Sparks Lawsuit

On June 19, 2026, a Chinese developer released the airllm 70B open-source model on GitHub, gaining 20,000 stars. Anthropic and other companies subsequently attempted to block the repository and filed a lawsuit.

AI Models 开源争议 Anthropic
547 06-21

Anthropic CEO Says India AI Summit Was "Extremely Chaotic," Modi Photo Session Sparks Political Debate

During the India AI Summit on June 19-20, 2026, Anthropic CEO Dario Amodei criticized the event as "extremely chaotic," specifically pointing to repeated adjustments for a group photo requested by Prime Minister Modi, which disrupted technical discussions and sparked political controversy.

Anthropic 印度AI峰会 AI国际合作
338 06-21

Qwen3 Max Main Score Plummets 19.2 Points, Code Execution Drops 31.2 Points in a Single Day

In the YZ Index's June 2026 test of 11 models, Qwen3 Max's main score dropped from 100 points yesterday to 80.82 points today, a decrease of 19.2 points, driven primarily by a sharp 31.2-point plunge in the code execution dimension.

Qwen3 Max Code Execution Smoke Test
488 06-21

Grok 4 Trend Up 19.8 Points Leads Smoke Weekly Report, Gemini Series Volatility Exceeds 28 Points

In the Smoke quick test from June 17 to 21, 2026, Grok 4 rose from 80.2 on the first day to 100 on the last day, a trend increase of 19.8 points, leading the weekly report. The Gemini series showed volatility exceeding 28 points.

Grok 4 Gemini 2.5 Pro Smoke 周趋势
521 06-21

Qwen3 Max Plunges 19.2 Points on Main Leaderboard; Four Models Score Perfect in Execution and Constraint

On 2026-06-21, the Smoke Lightweight Evaluation shows that four models — DeepSeek V4 Pro, Gemini 3.1 Pro, GPT-o3, and Grok 4 — all achieved 100 points in the main leaderboard, code execution, and material constraint scores, forming a perfect match between execution and constraint.

Qwen3 Max Code Execution Smoke Light Test
331 06-21

Dell AI Server Revenue Surges 757%, Order Backlog Reaches $51.3 Billion, Igniting Hardware Market Attention

Dell Technologies reported a massive 757% year-over-year increase in AI server revenue, with orders backlog reaching $51.3 billion, reflecting strong demand for AI infrastructure.

戴尔 AI Servers Earnings
323 06-21

Anthropic Mythos Model Emerges as Global AI Regulatory Policies Intensify

Anthropic released its new Mythos model, advancing reasoning and safety alignment, amid intensifying global AI regulation. The model faces external pressures from policy shifts and educational restrictions that may impact its deployment.

Anthropic Mythos AI Regulation
386 06-21

Karpathy Warns Developers: LLM Applications Go Beyond Prompt Engineering, Building Autonomous Systems Is the Right Path

Andrej Karpathy, former Chief Scientist at OpenAI, urges developers to move beyond prompt engineering and simple LLM calls. He emphasizes that building autonomous, self-improving systems is key to unlocking the full potential of large language models.

Karpathy LLM应用 AI开发
329 06-21

AI Infrastructure Decade-Long Construction Wave: Accelerated Full-Chain Deployment from Chips to Nuclear Energy

In recent years, the rapid development of AI has pushed global infrastructure construction into an unprecedented decade-long wave. Investors' attention has extended from single-chip fields to memory, data centers, power supply, nuclear energy, and even frontier areas such as robotics and quantum computing.

AI Infrastructure Data Centers 芯片
302 06-21

AI Stocks Account for Nearly 40% of S&P 500, Bubble Controversy Sparks Market Caution

AI-related stocks now represent 39% of the S&P 500, igniting widespread debate over a potential bubble. The concentrated rally, fueled by tech giants and interlocking investments, raises concerns about artificially created demand and valuation risks.

AI泡沫 S&P 500 Big Tech
295 06-21

GLM 5.2 Open-Source Breakthrough: Local AI Welcomes Its "ChatGPT Moment"

The release of GLM 5.2 marks a breakthrough for open-source local AI, offering a 1 million token context window and performance close to Opus 4.8, potentially accelerating the migration of AI from cloud to edge devices.

GLM 5.2 本地AI 开源模型
243 06-21

Local AI Agents and Offline Coding: Developer Community Heats Up Over Claude Code Practices

As AI technology evolves rapidly, local AI agents are becoming a focal point for developers. Recent discussions on X platform reveal how developers are leveraging tools like Claude Code to build local AI agents, enabling true offline coding and agent collaboration.

Local AI AI Agents Offline Coding
294 06-20

Accenture Stock Plunges 18%: How AI is Reshaping the Future of Consulting Industry

Accenture's stock plunged 18% after reporting earnings and lowering revenue guidance, highlighting how generative AI is rapidly replacing traditional IT consulting and outsourcing services.

埃森哲 AI影响 咨询行业
345 06-20

Open-Source GLM-5.2 Challenges Closed-Source Dominance: Coding Performance Nears Top Models, Igniting AI Community

Zhipu AI's newly released open-source GLM-5.2 model has drawn widespread attention for its coding capabilities, approaching those of top closed-source models and rekindling debates on AI openness.

GLM-5.2 open source LLM
316 06-20

Trump Administration's Ban Sparks Controversy: Anthropic Fable 5 Model Faces Risk of Shutdown

The Trump administration's ban or restrictions on Anthropic's latest AI models, Fable 5 and Mythos, have sparked an uproar in the tech world. Citing national security concerns, the policy requires the suspension or removal of these models, potentially reshaping the global AI competition landscape.

Anthropic AI regulation US ban
253 06-20

Claude Fable 5 and Mythos 5 Globally Removed on June 12 Amid Security Verification Demands and Privacy Controversy

Anthropic's Claude Fable 5 and Mythos 5 were globally removed on June 12, 2026 due to jailbreak vulnerability concerns and have not yet been restored. The incident also sparked controversy over new identity verification and biometric data collection requirements.

AI Models 安全政策 Anthropic
870 06-20

Anthropic Pauses Claude Agent SDK Token Billing Change Amid User Opposition

On May 13, 2026, Anthropic announced it would decouple Claude Agent SDK usage from subscription quotas, billing agent calls at standard API token prices starting June 15, with subscription fees serving only as monthly credit deductions. Following widespread opposition, the company paused the change around June 16.

Anthropic Claude AI计费
745 06-20

ERNIE Bot 4.5 Smoke Main Ranking Plunges 22.2 Points, Code Execution Halved to 50 Points

In the June 2026 YZ Index testing of 11 models, ERNIE Bot 4.5 Smoke's main ranking score dropped from 93.25 to 71.02, a single-day decline of 22.2 points.

ERNIE Bot 4.5 Code Execution Smoke测试
415 06-20

GPT-5.5 Smoke Mainboard Drops 20.5 Points, Code Execution Falls from 100 to 50

GPT-5.5’s mainboard score in today’s Smoke evaluation dropped 20.5 points, driven by a 50-point plunge in code execution, though analysis attributes the fluctuation to random question draws rather than model degradation.

GPT-5.5 Code Execution Smoke快测
348 06-20
20 21 22 23 24

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0