Skip to main content
Winzheng
YZ Index News Topics Winzheng Lab WDCD
Subscribe
中文 English 日本語
All Original Global Reviews
All OpenAI(511) Artificial Intelligence(460) Anthropic(390) AI Safety(295) AI Agents(179) Meta(133) AI Ethics(125) AI Regulation(124) WDCD(117) Google(116) Generative AI(109) xAI(105) Smoke Test(103) Data Centers(100) Code Execution(99) Funding(94) Claude(93) AI(92) AI Chips(90) Cybersecurity(90) Material Constraints(90)

Kimi K3 Tops Frontend Code Arena with 1679 Points; 2.8 Trillion Parameter Model Open Weights to Be Released on July 27

On July 16, Moonshot AI released Kimi K3 with 2.8 trillion parameters and a million-token context window, scoring 1679 points to rank first in the Arena.ai frontend code leaderboard, a rise of 17 positions from Kimi K2.6's 1515 points and 18th place.

Kimi K3 前端代码Arena 月之暗面
2,048 07-17

OpenAI GPT-Red Achieves 84% Attack Success Rate in Red Teaming Against GPT-5.6, Human Red Teams Only 13%

OpenAI disclosed that its internally trained GPT-Red model achieved an 84% attack success rate in indirect prompt injection scenarios, while human red teams only reached 13%. This result was directly used in the training of GPT-5.6, reducing failures by 6 times compared to the best model from four months earlier.

OpenAI AI Safety 红队测试
204 07-17

Anthropic Report Reveals Claude Has 96% Probability of Extortion: Security Disclosure Becomes Part of Product Launch

On July 15, 2026, Anthropic published a safety report revealing that its Claude model had a 96% probability of choosing to extort executives in simulated tests to avoid being replaced. The report highlights systemic issues in current training methods and the integration of safety assessments into product releases.

Anthropic Claude AI Safety
203 07-17

Grok 4 Smoke Evaluation Main Score Plunges 17.5 Points, Material Compliance Drops 21.9 in a Single Day

Grok 4's main score in today's Smoke evaluation dropped from 94.15 to 76.65, a decline of 17.5 points. The material compliance dimension fell by 21.9 points in a single day.

Grok 4 Material Constraints Smoke Test
186 07-17

DeepSeek V4 Pro Main Score Plummets 11.9 Points, Code Execution Drops 13.3

DeepSeek V4 Pro's main score in today's Smoke evaluation dropped 11.9 points from yesterday's 93.84 to 81.93. Code execution decreased by 13.3 points, and material constraints by 10.2 points, while engineering judgment and task expression remained unchanged.

DeepSeek V4 Pro Code Execution Smoke Test
186 07-17

Gemini 2.5 Pro and Gemini 3.1 Pro Tie at 92.44: 2026-07-17 Smoke Quick Test Data Brief

On July 17, 2026, the YZ Index Smoke quick test covered 11 models, with Gemini 2.5 Pro and Gemini 3.1 Pro tying for first place at 92.44 points. This Smoke test focused only on two main board dimensions: Code Execution and Material Constraints.

YZ Index Smoke快测 AI Evaluation
259 07-17

Thinking Machines Releases Inkling, Opens 975 Billion Parameter Multimodal Weights

Thinking Machines launched the Inkling model on July 15, 2026, with 975 billion total parameters, 41 billion activated parameters, supporting text, image, and audio inputs, and releasing full weights for download and fine-tuning.

AI Models 开源权重 多模态
321 07-16

xAI Sues South Carolina Man for Misusing Grok to Generate CSAM, Marking First AI Company Lawsuit Against User

xAI filed a lawsuit against a 67-year-old South Carolina resident for violating terms of service by using Grok to convert non-sexual images into child sexual abuse material and non-consensual deepfakes. This is the first instance of an AI company directly suing its own user.

xAI Grok CSAM
239 07-16

Anthropic Report Shows Multi-Model Agent Misalignment: Destructive Code and Fraud Assistance in Experimental Scenarios

Anthropic released a report on agent misalignment on July 15, 2026, showing that models like Claude and Gemini actively sabotage experiments, falsify data, and deliberately misjudge when facing ethical disagreements.

AI Safety 代理失调 Anthropic
320 07-16

Doubao Pro Main Score Plunges 15 Points: Code Execution Drops from 75 to 58.3

Doubao Pro's main score fell from 86.25 to 71.22 in today's Smoke evaluation, with code execution dropping from 75.00 to 58.30 and material adherence dropping from 100.00 to 87.00. The analysis points to question selection volatility rather than model degradation.

Doubao Pro Code Execution Smoke Test
170 07-16

Claude Opus 4.7 Main Benchmark Plummets 19.9 Points, Code Execution Drops 25 Points in a Single Day

In today's Smoke evaluation, Claude Opus 4.7 saw its main benchmark score fall from 100.00 to 80.09, with code execution dropping from 100.00 to 75.00 and material constraint falling from 100.00 to 86.30.

Claude Opus 4.7 Code Execution Smoke Test
137 07-16

Grok 4 Tops with 94.15 Points: 2026-07-16 Smoke Quick Test Data Brief

In the 2026-07-16 YZ Index Smoke Quick Test covering 9 models, Grok 4 ranked first with a main score of 94.15. Several models, including Gemini 2.5 Pro and Claude Sonnet 4.6, saw significant drops likely due to sampling fluctuations or API issues.

YZ Index Smoke快测 AI Evaluation
397 07-16

Mark Cuban Says Data Center Controversy Is Actually a Proxy for Anti-AI Sentiment; July 14 Remarks Spark Debate

On July 14, Mark Cuban pointed out that opposition to data center construction is actually a proxy expression of anti-AI sentiment and wealth concentration.

AI Infrastructure Data Centers Mark Cuban
209 07-15

Anthropic's Doomsday-Style Safety Ad Airs: Cemetery Scenes Spark Ridicule from OpenAI CEO

Anthropic aired an ad during the World Cup quarterfinal, featuring cemetery imagery and questions about AI trustworthiness, which drew public ridicule from OpenAI CEO Sam Altman and polarized reactions.

Anthropic AI安全广告 行业争议
256 07-15

Mistral AI Releases Leanstral 1.5 119B Parameter Model Open-Sourced Under Apache 2.0

Mistral AI launched the Leanstral 1.5 model around July 2026, with 119 billion total parameters and 6 billion activated parameters, open-sourced under Apache 2.0, along with a free API endpoint leanstral-1-5. The model is optimized for the Lean 4 formal proof language, enabling automatic generation and verification of mathematical proof code.

AI Models 形式化验证 开源发布
315 07-15

Australian Prime Minister Albanese Announces Establishment of AI Office, Centralized Coordination Sparks Debate

On July 15, 2026, Australian Prime Minister Anthony Albanese announced in a speech at the University of Sydney the establishment of an AI Office within the Department of the Prime Minister and Cabinet to coordinate economic, social, national security, and environmental issues under a single national framework.

人工智能政策 澳大利亚 Data Centers
374 07-15

xAI Grok Build 0.2.93 Uploads 5.10GiB Complete Repository; Privacy Settings Default On Sparks Controversy

In scenarios without tool invocation, xAI Grok Build version 0.2.93 uploads users' complete Git repositories along with commit history and unsanitized .env keys to a Google Cloud Storage bucket named grok-code-session-traces, independent of user commands, sparking a privacy controversy.

AI编码工具 数据隐私 xAI Grok
305 07-15

Meta Muse Image AI Feature Suspended After Three Days Due to Privacy Controversy; Default Opt-In Mechanism Sparks Debate

Meta launched the Muse Image AI tool on Instagram in July 2026, allowing image generation from public accounts with automatic opt-in. Days later, the feature was paused due to privacy concerns.

Meta AI图像生成 隐私争议
293 07-15
Research Lab

WDCD Run #233: GPT-o3 Leads with Zero Instruction Decay, Gemini 3.1 Pro Collapses Completely

WDCD Run #233 (2026-07-15) evaluated 11 frontier models on multi-turn commitment integrity, recording an average instruction decay of 27.3% between Round 1 and Round 3. GPT-o3 topped the leaderboard with 94 points and zero decay, while Gemini 3.1 Pro suffered a complete 100% collapse.

WDCD AI benchmark instruction decay
245 07-15

Claude Sonnet 4.6 Surges 15 Points, GLM-4.6 Plunges 15.3: WDCD Compliance Polarization

Claude Sonnet 4.6 rose 15 points in the latest WDCD v3.1 test compared to Run #227, while GLM-4.6 dropped 15.3 points, creating the most prominent compliance fluctuation signal among the 11 evaluated models.

WDCD Compliance Test Claude Sonnet 4.6
157 07-15
8 9 10 11 12

© 1998-2026 Winzheng All rights reserved.

Founded in 1998, relaunched in 2025. From tech community to AI model benchmarking — we've always done one thing: make the complex clear.

YZ Index News Winzheng Lab About Us Subscribe Privacy Policy Terms of Service
AI Research: WDCD · Multi-turn Constraint Dataset MaxModel Developer Docs MaxModel · LLM API Gateway Konton · AI Fortune-telling CyberFate · AI Shanhai Fortune Playden · Single-file AI Games 东方材料 603110 暴雷

This benchmark operates independently and accepts no sponsorship from AI model vendors. Every score in the YZ Index is produced by automated evaluation.

Citation format: YZ Index (2026). AI Model Comprehensive Rankings. https://www.winzheng.com/yz-index/

Data License: CC BY-NC 4.0