Doubao Pro Tops with 95.28 Points: 2026-09-11 Smoke Quick-Test Data Brief

On 2026-09-11, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first that day at 95.28. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.

Daily Ranking

RankingModelMain leaderboardCode executionMaterial constraintsIntegrity
#1Doubao Pro95.2810089.5pass
#2Gemini 3.1 Pro93.4810085.5pass
#3GPT-o393.4810085.5pass
#4GPT-5.590.8283.3100pass
#5Grok 488.6497.577.8pass
#6DeepSeek V4 Pro86.0983.389.5pass
#7Gemini 2.5 Pro81.537589.5pass
#8Claude Opus 4.779.17584.1pass
#9Claude Sonnet 4.674.8972.577.8pass
#10Qwen3 Max70.917565.9pass

Data Interpretation

In today's YZ Index Smoke quick test, Doubao Pro ranked first on the main leaderboard with 95.28; its combination of 100 in code execution and 89.5 in material constraints shows a balanced advantage. Gemini 3.1 Pro and GPT-o3 tied on the main leaderboard at 93.48, with both scoring 100 in code execution and 85.5 in material constraints, reflecting a structural emphasis on code execution. GPT-5.5's main leaderboard score of 90.82 relied on 100 in material constraints and 83.3 in code execution, forming a complementary profile. Leading models show clear divergence in how they combine strengths and weaknesses across the two dimensions.

Compared with the previous comparable run, GPT-o3 rose 15.4 points on the main leaderboard and 30.2 points in code execution; Gemini 2.5 Pro fell 7.2 points on the main leaderboard, fell 25 points in code execution, and rose 14.5 points in material constraints; GPT-5.5 rose 6.2 points on the main leaderboard and 11.3 points in code execution; Grok 4 fell 5.2 points on the main leaderboard, rose 8.8 points in code execution, and fell 22.2 points in material constraints; Claude Opus 4.7 rose 5.2 points on the main leaderboard, rose 22.4 points in code execution, and fell 15.9 points in material constraints. Among these shifts, the sharp drop of -22.2 points in Grok 4's material constraints, the sharp drop of -25 points in Gemini 2.5 Pro's code execution, and the sharp drop of -15.9 points in Claude Opus 4.7's material constraints constitute anomalous signals. They may stem from question-sampling volatility or may indicate genuine degradation, and need follow-up runs to verify.

GLM-4.6 had incomplete data due to API failure/timeout and did not participate in this period's ranking. As Smoke is a small-sample, single-day signal, the above observations reflect only that day's score structure and should not be used for long-term inferences for now.

Main Changes

  • GPT-o3: Main leaderboard up 15.4 points, code execution +30.2 points
  • Gemini 2.5 Pro: Main leaderboard down 7.2 points, code execution -25 points, material constraints +14.5 points
  • GPT-5.5: Main leaderboard up 6.2 points, code execution +11.3 points
  • Grok 4: Main leaderboard down 5.2 points, code execution +8.8 points, material constraints -22.2 points
  • Claude Opus 4.7: Main leaderboard up 5.2 points, code execution +22.4 points, material constraints -15.9 points

Signals to Watch

  • Grok 4: Material constraints plunge -22.2 points
  • Gemini 2.5 Pro: Code execution plunge -25 points
  • Claude Opus 4.7: Material constraints plunge -15.9 points
  • GLM-4.6: Incomplete data (missing several evaluation dimensions, including execution, judgment, integrity, and communication; API failure/timeout), has entered automatic rerun, and is not included in this period's ranking

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.


Data source: YZ Index | Run #318 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!