Doubao Pro Leads with 89.91 Points: 2026-09-09 Smoke Quick-Test Data Brief

On 2026-09-09, the YZ Index Smoke quick test covered 10 models, and Doubao Pro ranked first that day with 89.91 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers the two main benchmark dimensions of code execution and material constraints. The main benchmark formula is 0.55 × Code Execution + 0.45 × Material Constraints. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMain ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Doubao Pro89.9194.584.3pass
#2Grok 488.4293.582.2warn
#3GPT-o386.0494.575.7pass
#4GPT-5.584.1194.571.4pass
#5Claude Sonnet 4.681.8694.566.4pass
#6Claude Opus 4.780.9869.595pass
#7DeepSeek V4 Pro76.4869.585pass
#8Gemini 3.1 Pro74.2176.571.4pass
#9Gemini 2.5 Pro73.669.578.6pass
#10Qwen3 Max72.9769.577.2pass

Main Changes

  • GPT-5.5: Main score up 27.7 points; code execution +44.5 points; material constraints +7.1 points
  • GPT-o3: Main score up 24.8 points; code execution +44.5 points
  • Claude Opus 4.7: Main score up 22.2 points; code execution +44.5 points; material constraints -5 points
  • Qwen3 Max: Main score up 19.9 points; code execution +25.7 points; material constraints +12.9 points
  • Doubao Pro: Main score up 15.6 points; code execution +20.7 points; material constraints +9.3 points

Signals to Watch

  • Grok 4: Material constraints plunged -17.8 points
  • Claude Sonnet 4.6: Material constraints plunged -33.6 points
  • DeepSeek V4 Pro: Material constraints plunged -15 points
  • GLM-4.6: Data incomplete (several evaluation dimensions missing; API failure/timeout); auto-retry initiated; not ranked this round

When reading Smoke briefs like this, the focus should be on two questions: first, whether a given model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or material-constraint scores may stem from question sampling, or may be early signs of genuine degradation, and require follow-up runs for confirmation.


Data source: YZ Index | Run #315 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!