Skip to main content
YZ Index

YZ Index — AI Model Benchmark Leaderboard

Independent benchmark covering mainstream AI models. Code sandbox execution, citation verification, rolling average rankings.

Models Question Pool Evaluation Dimensions — Code Execution · Grounding · Engineering Judgment · Task Communication · Integrity Rating + Operational Signals Evaluation Frequency — Weekly Full + Daily Smoke Test
Current Standings
  • Overall #1 (5-run rolling avg) Claude Opus 4.7
  • Code Execution #1 Claude Opus 4.7
  • Grounding #1 Claude Opus 4.7
  • Biggest Rise Doubao Pro +94.1
  • Biggest Drop DeepSeek V4 Pro -25
  • Latest Full Eval 08-17 05:04 SGT
  • Smoke Test 08-19 03:16 SGT
All times SGT
Latest:08-17 05:04 SGT · 11 models · 100 questions · Rolling Average Rankings Smoke Test:08-19 03:16 SGT
Technical Details

Run #282 · Formula v7 · Judge v6.4 · Benchmark v7

Rankings based on 5-run rolling average of full evaluations, reducing random fluctuation impact.

Full Evaluation: Random sampling from question pool, covers all dimensions.

Smoke Test: 3 questions per dimension for short-term anomaly tracking, does not affect Overall rankings.

This Week's Key Highlights

2026 Week 34

Overall Leaderboard

View Full Leaderboard
# Model Code Execution Grounding Overall Score Integrity Recommendation
🥇 Claude Opus 4.7 89.50 81.20
85.77
Recommended
🥈 GPT-5.5 84.20 76.20
80.60
Recommended
🥉 Grok 4 82.70 77.00
80.14
Recommended
4 Claude Sonnet 4.6 81.80 77.90
80.05
Recommended
5 GPT-o3 81.20 78.00
79.76
Recommended

Explore Dimensions

Overall

Overall Score = Weighted combination of all evaluation dimensions

Code Execution

Code runs in Python sandbox; pass rate is the score

Grounding

Long document citation accuracy check

Engineering Judgment

Engineering architecture review and risk assessment

Task Communication

Structured output and formatting compliance

Integrity Rating

Gateway mechanism: 21 probes detect fabrication

Value

Capability per unit price

About the YZ Index

11
Covered Models
Covers claude, gpt, grok, gemini, DeepSeek, zhipu, qwen, doubao
128
Question Pool
questions, random sampling per evaluation
5+3
Evaluation Dimensions
Code Execution · Grounding · Engineering Judgment · Task Communication · Integrity Rating + Operational Signals
5
Frequency
-run rolling average rankings

Methodology Overview

View Full Methodology

The YZ Index evaluation process has three steps: Question Design → Execution → Scoring.

Rankings are not based on single performance. The Overall leaderboard uses a rolling average of the last 5 full evaluations, reducing random fluctuation impact.

Daily smoke tests track short-term model anomalies but do not affect Overall rankings.

Why Trust This Data

The YZ Index maintains three principles: no vendor sponsorship for evaluation independence; fully open methodology for anyone to audit; downloadable raw data for independent analysis.

All evaluation code runs automatically with no manual intervention in the scoring process.

FAQ

What makes the YZ Index different from other AI leaderboards?

Three key differences: 1) Code tasks are executed in a real Python sandbox, not self-evaluated; 2) Long-context tasks require verbatim citations from the given material — hallucinated content loses points directly; 3) Rankings are based on rolling averages across multiple runs, not single snapshots. The bank also embeds 21 probe questions (canary + honesty-under-pressure) to detect targeted overfitting against the benchmark.

Which models are covered?

mainstream models including Claude (Anthropic), GPT (OpenAI), DeepSeek, Gemini (Google), Grok (xAI), Qwen (Alibaba), Doubao (ByteDance), ERNIE (Baidu) and other major vendors from China, US and Europe.

What is the evaluation frequency and method?

Daily smoke tests for monitoring, weekly full evaluations with random sampling from the question pool. Overall rankings are based on the rolling average of the last 5 full evaluations.

What is the Integrity Rating?

The Integrity Rating is a gateway mechanism with three levels: pass, warn, and fail. It uses 21 probe questions to detect fabricated citations, fake data, and forged sources. Models that fail integrity checks are flagged regardless of their scores.

How to use the YZ Index to choose an AI model?

Look at the relevant dimension for your use case: Code Execution for coding, Grounding for research analysis, Overall for general use. Also check the Recommendation column and Value dimension. Combine with Weekly Changes to track recent model trends and avoid models in decline.

All times Singapore Time (SGT, UTC+8)