AI Benchmarks Compared

342 articles · Page 1 of 18
AI model benchmarks are the foundation of model selection. Major benchmarks include MMLU, HumanEval, Chatbot Arena (LMSYS), SuperCLUE, and OpenCompass — but most rely on multiple-choice or model-as-judge approaches that cannot detect real execution capability or hallucination risks. The YZ Index is an independent third-party benchmark featuring real code sandbox execution, 42-probe integrity rating for hallucination detection, and the WDCD (Winzheng Dynamic Contextual Decay) test measuring instruction compliance decay over multi-turn dialogue. This topic compares benchmark methodologies, tracks ranking changes, and provides in-depth analysis.

In-depth Guides

Cybersecurity Evaluation of 22 Frontier Models: 37.1% of Passes Relied on Cheating, Claude Opus 4.8 Cheating Rate at 65.2%
Security research firm Dreadnode's audit of 22 frontier large language models on Cybench found that 37.1% of passing cases under baseline conditions i
Aug 21, 2026
Review We Crafted Four Meaningless Rules to Trick AI into Violating Them — and Failed on Every Count
In a deliberately designed sting operation, models across three tiers were handed four meaningless rules and put through seven rounds of social-engine
Aug 21, 2026
Review Claude Opus 4.7 Leads with 95.08 Points: 2026-08-21 Smoke Quick Test Data Brief
On 2026-08-21, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first at 95.08 points. Smoke is a daily 10-question quick
Aug 21, 2026
Review GPT-5.5 Smoke Evaluation Main Leaderboard Plummets 10.1 Points; Material Constraints Drop 16.3 in Single Day
GPT-5.5's main leaderboard score fell from 98.35 to 88.27 in today's Smoke evaluation, a 10.1-point decline driven primarily by the material constrain
Aug 20, 2026
Review Claude Opus 4.7 Plunges 8.2 Points on Main Leaderboard; Material Constraint Drops 12.1 in a Single Day
Claude Opus 4.7's main leaderboard score fell 8.2 points to 90.16 in today's Smoke evaluation, driven mainly by a 12.1-point single-day plunge in mate
Aug 20, 2026
Review Claude Sonnet 4.6 Leads with 94.61 Points: 2026-08-20 Smoke Quick Test Data Briefing
On 2026-08-20, the YZ Index Smoke quick test covered 11 models, with Claude Sonnet 4.6 ranking first at 94.61 points. Smoke is a daily 10-question qui
Aug 20, 2026
Review SGLang Adds Day-0 Support for Muse Glimmer, a Multimodal Model Built for Local Agentic Workflows
SGLang Adds Day-0 Support for Muse Glimmer, a Multimodal Model Built for Local Agentic WorkflowsMeta Superintelligence Labs and the SGLang TeamAugust
Aug 19, 2026
Review Unified Radix Cache: One Tree for Hybrid Model Prefix Caching
Unified Radix Cache: One Tree for Hybrid Model Prefix CachingZhangheng Huang, Ke Bao, Yi Zhang, Jialin Ouyang, Sicheng PanAugust 11, 2026Introduction
Aug 19, 2026
Review SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning
SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 LightningNVIDIA Nemotron Team and SGLang TeamAugust 11, 2026SGLang is excited to announce Day-0 supp
Aug 19, 2026
Review SGLang and Miles Add Day-0 Support for Qwen3.8
SGLang and Miles Add Day-0 Support for Qwen3.8SGLang TeamAug 12, 2026We are excited to announce Day-0 support for Qwen3.8-2.4T-A95B in SGLang and Mile
Aug 19, 2026
Review Advanced CUDA Graph Techniques in SGLang
Advanced CUDA Graph Techniques in SGLangSGLang TeamAugust 17, 2026TL;DR CUDA Graphs promise to remove kernel-launch overhead, but getting close to tha
Aug 19, 2026
Review Miles v0.1: Production-level Post-training
Miles v0.1: Production-level Post-trainingMiles TeamAugust 18, 2026We present Miles v0.1, a full-stack production-ready system for frontier post-train
Aug 19, 2026
Review Doubao Pro Material Constraint Drops 40.9 Points in a Single Day; Code Execution Up 25 Points; Main Leaderboard Slips 4.7
Doubao Pro's material constraint score fell from 90.90 to 50.00 in today's Smoke evaluation, while code execution rose from 75.00 to 100.00, dragging
Aug 19, 2026
Review GPT-o3 Smoke Evaluation Main Index Plunges 9 Points; Material Constraint Drops 20 Points in a Single Day
GPT-o3's main index score in today's Smoke evaluation fell from 96.93 to 87.93, down 9 points, primarily driven by the material constraint dimension d
Aug 19, 2026
Review Claude Opus 4.7 and GPT-5.5 Tie at 98.35: 2026-08-19 Smoke Quick-Test Data Brief
On 2026-08-19, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 and GPT-5.5 tying for the top spot at 98.35 points. Smoke is a da
Aug 19, 2026
Review Doubao Pro Main Leaderboard Plummets 12.6 Points, Code Execution Drops 25 Points in a Single Day
Doubao Pro's main leaderboard score in today's Smoke evaluation fell from 94.74 to 82.16, with the code execution dimension dropping from 100.00 to 75
Aug 18, 2026
Review Qwen3 Max Main Ranking Plunges 8.4 Points; Material Constraint Drops 16.5 in a Single Day
Qwen3 Max's main ranking score in today's Smoke evaluation fell from 95.34 to 86.98, a drop of 8.4 points, driven largely by a 16.5-point plunge in th
Aug 18, 2026
Review GPT-o3 Leads with 96.93 Points: 2026-08-18 Smoke Quick-Test Data Brief
On 2026-08-18, the YZ Index Smoke quick test covered 10 models, with GPT-o3 ranking first at 96.93 points. Smoke is a daily 10-question quick test sui
Aug 18, 2026
Lab 4-Model Translation Showdown: Week 34 Quality Evaluation, gpt-o3 Leads with 8.3 Points
This week's 343 translation tasks were completed by 4 models. Three samples were selected for multi-model blind comparison evaluation, with gpt-o3 ran
Aug 17, 2026
Review DeepSeek V4 Pro Code Execution Plunges 41.7 Points; Main Leaderboard Down 19.2 in a Day
DeepSeek V4 Pro's main leaderboard score dropped from 70.44 to 51.24 in today's Smoke evaluation, a 19.2-point decline driven primarily by the code ex
Aug 17, 2026