YZ Index — AI Model Benchmarks, News & Research
Editor's Pick
The U.S. is building barriers around drones and robots, but China has scale to get around them
The U.S. is shutting out more foreign-made drones and robots. China’s scale means the global competition may simply move elsewhere.
2026-08-31 11:23
EU Designates ChatGPT as Very Large Online Search Engine, OpenAI Faces Strictest DSA Compliance Obligations
On August 31, 2026, the European Commission designated ChatGPT as a Very Large O
Meeting notetaker Circleback adds a free tier to attract more customers
Circleback is also introducing new pricing plans starting from $14 per month
Overall Top 5
Full Rankings →
#1
Grok 4 81.6
▲7.6
·
#2
Claude Opus 4.7 80.9
▼1.7
·
#3
GPT-o3 80.2
▲2.1
·
#4
豆包 Pro 79.8
▲1.5
·
#5
GPT-5.5 76.1
·
#6
Claude Sonnet 4.6 75.2
▼2.3
·
#7
Qwen3 Max 72.9
▼1.1
·
#8
Gemini 2.5 Pro 71.9
▲0.6
·
#9
DeepSeek V4 Pro 70.3
▲2.8
·
#10
Gemini 3.1 Pro 66.6
▼3.4
·
#11
GLM-4.6 66.5
▲13.2
·
▲ Claude Opus 4.7 +15 · ▼ GLM-4.6 -30.8
·
#1
Grok 4 81.6
▲7.6
·
#2
Claude Opus 4.7 80.9
▼1.7
·
#3
GPT-o3 80.2
▲2.1
·
#4
豆包 Pro 79.8
▲1.5
·
#5
GPT-5.5 76.1
·
#6
Claude Sonnet 4.6 75.2
▼2.3
·
#7
Qwen3 Max 72.9
▼1.1
·
#8
Gemini 2.5 Pro 71.9
▲0.6
·
#9
DeepSeek V4 Pro 70.3
▲2.8
·
#10
Gemini 3.1 Pro 66.6
▼3.4
·
#11
GLM-4.6 66.5
▲13.2
·
▲ Claude Opus 4.7 +15 · ▼ GLM-4.6 -30.8
·
YZ Index · Weekly real-sandbox evaluation of 11 mainstream models · Zero vendor sponsorship · Auditable scoring Methodology →
Latest News
View All News →
ChatGPT and Reddit now face EU's toughest online safety rules
Explosive growth comes with a new regulatory burden in the European Union.
EU Designates ChatGPT as Very Large Online Search Engine, OpenAI Faces Strictest DSA Compliance Obligations
On August 31, 2026, the European Commission designated ChatGPT as a Very Large Online Search Engine under the Digital Se
Meeting notetaker Circleback adds a free tier to attract more customers
Circleback is also introducing new pricing plans starting from $14 per month
SB Energy Grants OpenAI $5.5 Billion in Warrants to Lock In 10GW Stargate Data Center in Ohio
SB Energy granted OpenAI warrants valued at $5.5 billion in exchange for a 20-year lease at a 10GW Stargate data center
OpenAI Secretly Purchases Tens of Thousands of Mac Minis: How Apple's M Chips Are Tearing Open the AI Computing Landscape
According to a report by The Information published on August 31, 2026, OpenAI has quietly assembled a massive fleet of A
US Commerce Department Plans to Ban Remote Computing Power Leasing: From Chip Controls to Compute Controls, China's AI Faces New Hurdles
The US Commerce Department is drafting an unprecedented export control rule to bar Chinese AI firms from remotely access
You Know Who Really Hates AI? Insurance Claims Adjusters
Of the Glassdoor reviews from claims adjusters that mentioned AI, a staggering 98 percent were negative. “AI is just a t
Pocket's AI made my game ideas real. Now Meta controls the results.
Interactive mobile "gizmos" are easy to make, hard to share outside Meta's platform.
Three US States Tighten AI Data Center Approvals in Succession; 474 GW of Power Applications Triggers Infrastructure Governance Crisis
Three US states have moved within six weeks to restrict AI data center expansion amid a 474 GW backlog of grid interconn
China's Ten Departments Jointly Implement AI "Review-Before-Creation": What the World's First Mandatory Pre-Research Ethics Review Mechanism Means
On April 2, 2026, ten Chinese government departments jointly issued the "AI Science and Technology Ethics Review and Ser
EU AI Office Sends First RFIs to OpenAI and Others on August 29, Formally Launching Mandatory Enforcement Phase
On August 29, 2026, the EU AI Office issued its first information requests to general-purpose AI model providers includi
DeepSeek Raises 50 Billion Yuan at $74 Billion Valuation, Begins Preparations for 2027 STAR Market IPO
DeepSeek is close to completing a funding round of approximately 50 billion yuan at a valuation of about 500 billion yua
Reviews
View All →Grok 4 Leads with 89.15: 2026-08-31 Smoke Quick Test Data Brief
On 2026-08-31, the YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first at 89.15 points. Smoke is a da
Gemini 3.1 Pro Main Score Plunges 10.8 Points; Material Constraint Drops 14.4 in a Single Day
Gemini 3.1 Pro's main score in today's Smoke evaluation fell from 96.60 to 85.83, a drop of 10.8 points, driven primaril
Grok 4 Main Ranking Plunges 8.8 Points: Code Execution Drops from 100 to 90.7
Grok 4's main ranking score in today's Smoke evaluation fell from 96.99 to 88.14, a single-day decline of 8.8 points, wi
WDCD Compliance
What it tests: whether AI holds your original instructions across multi-turn dialogue
#1
Grok 4
96.3
#2
GPT-o3
95.2
#3
GLM-4.6
93.7
#4
DeepSeek V4 Pro
92.2
#5
Claude Sonnet 4.6
91.6
#6
Gemini 3.1 Pro
88.2
#7
Claude Opus 4.7
85.8
View full compliance rankings →
Research Lab
4-Model Translation Showdown: Week 36 Quality Assessment, claude-sonnet-4.6 Leads with 9 Points
This week, 405 translation tasks were completed by 4 models. In a sampled blind comparison of 3 arti
WDCD Run #296: Zero Instruction Decay Across All 11 Models, Grok 4 Leads at 96.3
WDCD Run #296 (2026-08-26) recorded 0% average commitment decay across 11 tested models, with Grok 4
Translation Showdown of 4 Major Models: Week 35 Quality Review — gpt-o3 Leads with 8.3 Points
This week, 358 translation tasks were completed by 4 models. Three articles were sampled for multi-m