Doubao Pro Tops with 83.94 Points: 2026-09-12 Smoke Quick Test Data Brief

On 2026-09-12, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first that day at 83.94 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Doubao Pro83.9410064.3pass
#2Qwen3 Max83.599.264.3pass
#3Claude Opus 4.782.299764.3pass
#4GPT-5.582.299764.3pass
#5GPT-o381.0894.864.3pass
#6Grok 476.5896.252.6pass
#7Claude Sonnet 4.676.295.552.6pass
#8Gemini 2.5 Pro76.1173.579.3pass
#9DeepSeek V4 Pro63.277252.6pass
#10Gemini 3.1 Pro58.9154.564.3pass

Data Interpretation

In today's Smoke quick test, leading models showed a concentrated pattern in the pairing of code execution and material constraints. Doubao Pro scored 83.94 on the main leaderboard, with 100 in code execution and 64.3 in material constraints; Qwen3 Max scored 83.5 on the main leaderboard, with 99.2 in code execution and 64.3 in material constraints; Claude Opus 4.7 and GPT-5.5 both scored 82.29 on the main leaderboard, with 97 in code execution and 64.3 in material constraints; GPT-o3 scored 81.08 on the main leaderboard, with 94.8 in code execution and 64.3 in material constraints. These models had relatively high code execution scores while material constraints remained at 64.3, forming a structure of strong code execution paired with moderate material constraints.

Gemini 2.5 Pro scored 76.11 on the main leaderboard, with 73.5 in code execution and 79.3 in material constraints; its material constraints were relatively stronger. Among notable changes, Gemini 3.1 Pro fell 34.6 points on the main leaderboard, 45.5 points in code execution, and 21.2 points in material constraints; DeepSeek V4 Pro fell 22.8 points on the main leaderboard, 11.3 points in code execution, and 36.9 points in material constraints; Qwen3 Max rose 12.6 points on the main leaderboard and 24.2 points in code execution; GPT-o3 fell 12.4 points on the main leaderboard, 5.2 points in code execution, and 21.2 points in material constraints; Grok 4 fell 12.1 points on the main leaderboard and 25.2 points in material constraints. These signals may come from question sampling fluctuation or may be real degradation, and need follow-up runs for review. Smoke is a small-sample single-day signal, so interpretation should remain restrained.

Major Changes

  • Gemini 3.1 Pro: main leaderboard down 34.6 points, code execution -45.5 points, material constraints -21.2 points
  • DeepSeek V4 Pro: main leaderboard down 22.8 points, code execution -11.3 points, material constraints -36.9 points
  • Qwen3 Max: main leaderboard up 12.6 points, code execution +24.2 points
  • GPT-o3: main leaderboard down 12.4 points, code execution -5.2 points, material constraints -21.2 points
  • Grok 4: main leaderboard down 12.1 points, material constraints -25.2 points

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this kind of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be early signals of real degradation, and require follow-up runs for review.


Data source: YZ Index | Run #319 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!