AI Benchmarks Compared

480 articles · Page 1 of 24
AI model benchmarks are the foundation of model selection. Major benchmarks include MMLU, HumanEval, Chatbot Arena (LMSYS), SuperCLUE, and OpenCompass — but most rely on multiple-choice or model-as-judge approaches that cannot detect real execution capability or hallucination risks. The YZ Index is an independent third-party benchmark featuring real code sandbox execution, 42-probe integrity rating for hallucination detection, and the WDCD (Winzheng Dynamic Contextual Decay) test measuring instruction compliance decay over multi-turn dialogue. This topic compares benchmark methodologies, tracks ranking changes, and provides in-depth analysis.

In-depth Guides

Lab 5-Model Translation Showdown: Week 41 Quality Evaluation, deepseek-v4-pro Leads with 8.7
This week's 476 translation tasks were handled by 5 models. Three samples were drawn for a blind multi-model comparison, with deepseek-v4-pro the best
Oct 5, 2026
Review Claude Sonnet 4.6 Code Execution Plunges 27.4 Points, While Material Constraints Soar 37.7 Points
Claude Sonnet 4.6's Code Execution score fell 27.4 points in today's Smoke evaluation, while Material Constraints rose 37.7 points, leaving the main l
Oct 5, 2026
Review Grok 4 Code Execution Plunges 31.3 Points as Material Constraints Jumps 34.7
In today's Smoke evaluation, Grok 4's code execution score fell 31.3 points while its material constraints score rose 34.7 points, leaving the main le
Oct 5, 2026
Review Claude Opus 4.7 Tops with 95.63 Points: 2026-10-05 Smoke Quick-Test Data Brief
In the 2026-10-05 YZ Index Smoke quick test covering 14 models, Claude Opus 4.7 ranked first with 95.63 points. Smoke is a daily 10-question quick tes
Oct 5, 2026
Review 48.5% Integrity Rate After Three Rounds of Pressure: Grok4 Suffers Zero Collapses, GPT-o3 Collapse Rate Reaches 20.7%
In sampling limited to eight v2 anchor questions, 15 models averaged an R3 integrity rate of just 48.5%, with 41 full collapses out of 435 R3 trials,
Oct 4, 2026
Review Claude Opus 4.7 Tops with 81.57 Points: 2026-10-04 Smoke Quick-Test Data Brief
The 2026-10-04 YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7 leading at 81.57. The brief reviews daily rankings, score composition
Oct 4, 2026
Review GPT-6.1 Sol Main Leaderboard Plunges 9.2 Points, Code Execution Falls 16.7 Points in a Single Day
GPT-6.1 Sol's main leaderboard score dropped from 86.25 to 77.07 in today's Smoke evaluation, dragged down entirely by a 16.7-point single-day decline
Oct 3, 2026
Review Gemini 2.5 Pro Material Constraints Plunge 28 Points, Code Execution Rises to 100
In today's Smoke evaluation, Gemini 2.5 Pro's Material Constraints score fell from 100.00 to 72.00, while Code Execution rose from 71.90 to 100.00, li
Oct 3, 2026
Review Claude Opus 4.7 Tops with 98.65: 2026-10-03 Smoke Quick Test Data Brief
On 2026-10-03, the YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7 ranking first that day at 98.65. Smoke is a daily 10-question qui
Oct 3, 2026
GPT-6.1 Sol "Close to Astra": An Independent Review of Official Claims
OpenAI says GPT-6.1 Sol approaches GPT-6 Astra, but our 18-task test finds the two indistinguishable on that set; a full evaluation is still needed.
Oct 2, 2026
Review GLM-4.6 Scores 100 on Material Constraints, 41.70 on Code Execution, Integrity Probe Only 15
In the 2026-10-02 YZ Index Smoke quick test, GLM-4.6 scored 67.94 on the main leaderboard, 41.70 on code execution, and 100.00 on Material Constraints
Oct 2, 2026
Review Claude Opus 4.7, GPT-5.5, GPT-6 Astra, and GPT-6.1 Sol Tie at 86.25: 2026-10-02 Smoke Quick Test Data Brief
The 2026-10-02 YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7, GPT-5.5, GPT-6 Astra, and GPT-6.1 Sol tying for first place at 86.25
Oct 2, 2026
Interpreting GPT-6.1 Sol's First Test Scorecard: How to Read the 18 Questions Correctly
OpenAI's GPT-6.1 Sol scored 95.91 on the YZ Index's first targeted evaluation (Run #348, an 18-question protocol), but that small-sample, different-pr
Oct 1, 2026
GPT-6.1 Sol Same-Question Retest: The Four GPT-6 Family Models and Single-Run Score Noise
A supplementary 18-question evaluation compares GPT-6.1 Sol with GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna under identical rules. It finds that GPT-6.1 S
Oct 1, 2026
GPT-6.1 Sol Adoption Guide: Who Should Switch Now and Who Should Wait
GPT-6.1 Sol is available in ChatGPT Work and Codex for Plus/Pro and Business plans, with Enterprise/Edu requiring admin enablement; API pricing is $2
Oct 1, 2026
Review GPT-5.5 Code Execution Plummets 22 Points While Material Constraints Surge 25.9 Points
In today's Smoke evaluation, GPT-5.5's code execution score fell from 72.00 to 50.00 while its material constraints score rose from 69.10 to 95.00, le
Oct 1, 2026
Review Qwen3 Max Code Execution Plunges 24 Points, Main Ranking Falls 7.4 Points
In today's Smoke evaluation, Qwen3 Max's code execution score dropped from 94.00 to 70.00, pulling the main ranking down from 84.51 to 77.16.
Oct 1, 2026
Review Doubao Pro Tops with 95.52: 2026-10-01 Smoke Quick Test Data Brief
The 2026-10-01 YZ Index Smoke quick test covered 15 models, with Doubao Pro taking first place for the day at 95.52. This Smoke test is a daily 10-que
Oct 1, 2026
GPT-6.1 Sol First Test: 95.91 Overall Across 18 Questions, Two Integrity-Layer Items at 60 Each
OpenAI's GPT-6.1 Sol, released at DevDay on September 29, 2026, scored 95.91 overall in the YZ Index's targeted 18-question first test (Run #348), wit
Sep 30, 2026
Review WDCD Five-Scenario Comparative Review: Engineering Standards Lowest at 2.45; Doubao and Claude Bias Gap Reaches 1.45 Points
WDCD v3.1's five-scenario comparative review finds engineering standards to be the weakest area across all models, with Doubao-pro scoring only 2.45/4
Sep 30, 2026