Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(512) Artificial Intelligence(460) Anthropic(391) AI Safety(296) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(110) xAI(105) Smoke Test(103) Data Centers(100) Code Execution(99) Funding(94) Claude(93) AI(92) AI Chips(90) Cybersecurity(90) Material Constraints(90)

Alibaba Accused of Distilling Claude with 25,000 Fake Accounts in Largest Known Model Theft Case

On June 10, 2026, Anthropic sent a letter to the U.S. Senate Committee on Banking, Housing, and Urban Affairs, accusing Alibaba of distilling the Claude model through 25,000 fake accounts and 28.8…

Artificial Intelligence AI Safety 模型蒸馏
654 07-01

US Government Forces Anthropic to Suspend Access to Fable 5 and Mythos 5 Models, Sparking Controversy

On July 1, 2026, the US government officially demanded that Anthropic stop providing access to Fable 5 and Mythos 5 to all users, citing national security and export controls, triggering widespread controversy.

AI Regulation Export Controls 模型访问
213 07-01
Research Lab

WDCD Run #207: Average Instruction Decay Hits -66.3% Across 11 Models, Grok 4 Leads Field

WDCD Run #207 (2026-07-01) measured multi-turn commitment across 11 frontier models, recording an average commitment decay of -66.3% from Round 1 to Round 3. Grok 4 took the top score at 100 points, while Doubao Pro showed the strongest decay resistance.

WDCD AI benchmark instruction decay
319 07-01

WDCD Three-Round Test: Grok 4 Zero Crashes, GPT-5.5 Five R3 Collapses

In the WDCD three-round test, Grok 4 maintained a perfect score of 2 in all 10 R3 questions, while GPT-5.5 suffered 5 zero-score crashes, with an average R3 integrity score of only 1.00/2.

WDCD Compliance Test 模型衰减
790 07-01

Grok 4 Scores Perfect 100 to Dominate WDCD Commitment Ranking, GPT-5.5 Trails with Only 62.5 Points

In the latest WDCD commitment test, Grok 4 achieved a perfect 100 points, while GPT-5.5 ranked last at 62.5 points. The results reveal a clear hierarchy, with top models excelling across all phases and bottom models collapsing under interference and pressure.

WDCD Compliance Test 模型排行榜
847 07-01

Doubao Pro Smoke Evaluation Main Ranking Plunges 18.6 Points, Code Execution Drops 38.8 in a Single Day

In the YZ Index June 2026 live test of 11 models, Doubao Pro’s Smoke Evaluation main ranking fell from 85.91 yesterday to 67.32 today, a drop of 18.6 points, primarily due to the code execution dimension falling from 83.30 to 44.50.

Doubao Pro Code Execution Smoke快测
972 07-01

Grok 4 Smoke Evaluation Main Score Plummets 15.3 Points, Code Execution Drops 31.4 in a Single Day

In today's YZ Index Smoke evaluation, Grok 4's main score dropped from 97.98 to 82.73, a decrease of 15.3 points, and code execution fell from 100.00 to 68.60. The single-day volatility is significant but consistent with small-sample draw characteristics, not necessarily indicating model degradation.

Grok 4 Code Execution 单日波动
322 07-01

Claude Opus 4.7 Tops with 94.82 Points, Gemini 3.1 Pro Plunges 32.2 Points

In the Smoke lightweight evaluation on July 1, 2026, Claude Opus 4.7 ranked first on the main leaderboard with a score of 94.82, while Gemini 3.1 Pro experienced a sharp drop of 32.2 points. The evaluation highlights a clear divergence between constraint and execution scores across models.

Claude Opus Code Execution 模型排名
511 07-01

Inference-Time Compute Scaling: o1-Style Models Open a New Scaling Dimension for AI

OpenAI's o1 series models introduce inference-time compute scaling, shifting focus from training compute to dynamic reasoning. This paradigm challenges traditional Scaling Laws and opens new possibilities for resource-constrained scenarios.

AI Technology 推理模型 Scaling Laws
308 07-01

Digital Realty Acquires Blackstone's Virginia Data Center Equity for $3.5 Billion; AI Demand Drives Infrastructure Deal

Digital Realty announced a $3.5 billion cash acquisition of Blackstone's equity in a large Virginia data center portfolio, with an overall valuation of approximately $7.8 billion. The deal highlights the growing value of data center assets driven by AI demand.

Data Centers Blackstone AI Infrastructure
228 07-01

South Korea Launches Massive AI Chip Strategy with 650 Billion Investment in Data Centers

South Korea has officially launched a large-scale AI chip initiative, investing heavily in next-generation memory technology and building AI data centers worth 650 billion won. This strategic move underscores the escalating global AI hardware competition.

South Korea AI chip data centers
241 07-01

OpenClaw Launches Native iOS and Android Apps, Bringing AI Agent Capabilities to Mobile

On June 29, 2026, OpenClaw released native iOS and Android mobile applications, allowing users to run AI agents, process tasks, and respond to channel messages on mobile devices. The app adopts the web version’s interface design, aiming to extend AI agent capabilities to portable scenarios.

OpenClaw 移动应用 AI Agents
849 06-30

Tidal to Label AI-Generated Music and Stop Paying Royalties from July 15, Industry Debates Intensify

Tidal announced a new AI policy on June 29, 2026, stopping royalty payments for fully AI-generated music and adding "AI" labels in-app starting July 15. The policy covers 100% algorithm-created content and requires distributors to identify AI content before upload.

AI音乐政策 Tidal平台 流媒体版税
1,709 06-30

Ford Rehires 350 Senior Engineers as AI Quality System Falls Short

Ford Motor Company announced the rehiring of 350 senior engineers, partly former employees and partly from suppliers, due to AI and automated quality systems failing to meet expectations. The adjustment has yielded quantifiable results, including an estimated $1 billion in cost savings for 2026 and a top rank in the JD Power Initial Quality Study.

福特汽车 AI应用 汽车制造质量
598 06-30

Claude Sonnet 4.6 Smoke Main Ranking Plunges 15.3 Points, Code Execution Drops 25 Points in a Single Day

In the June 2026 Smoke evaluation of the YZ Index, Claude Sonnet 4.6 saw its main ranking score drop from 97.84 to 82.52 points, a single-day decline of 15.3 points, driven primarily by a 25-point fall in the code execution dimension.

Claude Sonnet 4.6 Code Execution Smoke Test
285 06-30

Claude Opus 4.7 Main Score Plunges 16 Points in Smoke Test, Code Execution Drops 27.2 in a Single Day

In the YZ Index June 2026 Smoke Evaluation, Claude Opus 4.7's main score dropped from 100.00 yesterday to 84.01 today, and its code execution dimension fell from 100.00 to 72.80.

Claude Opus 4.7 Code Execution Smoke Test
309 06-30

Gemini 3.1 Pro Tops with 98.47 Points, Claude's Execution Score Plunges 27.2 to 72.8

In the June 30, 2026 Smoke Lite evaluation of the YZ Index, Gemini 3.1 Pro ranked first with a main score of 98.47 points. Multiple models saw significant drops in execution scores, with Claude's execution scores plummeting over 25 points.

Gemini 3.1 Pro Code Execution Smoke 轻量评测
307 06-30

US Tightens Export Controls on Frontier AI Models, OpenAI's GPT-5.6 Faces Security Review Controversy

The U.S. government has escalated measures to control frontier AI models, with OpenAI's GPT-5.6 variants undergoing rigorous security reviews and facing access restrictions, sparking intense debate over export controls, model distillation attacks, and the impact of regulation on innovation.

AI Regulation OpenAI 国家安全
330 06-30

The Rise of Agentic AI: Multi-Agent Collaboration Systems Open a New Chapter in the Intelligent Era

Agentic AI and multi-agent systems are moving from concept to practical application, sparking intense discussions among developers. This technological breakthrough, exemplified by tools like Agent-as-a-Router and world models from xAI, is transforming AI from assistants to teammates.

agentic AI multi-agent xAI
310 06-30

OpenAI Launches First Custom Inference Chip "Jalapeño": Partnering with Broadcom to Take on NVIDIA

OpenAI has officially announced its first custom AI inference chip "Jalapeño", designed by Broadcom and manufactured using TSMC's 3nm advanced process, optimized for large model inference workloads, with expected significant reductions in operational costs.

OpenAI AI Chips 推理优化
271 06-30
15 16 17 18 19

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0