Doubao Pro Leads with 96.7 Points: 2026-08-02 Smoke Quick Test Data Briefing

On August 2, 2026, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first for the day at 96.7 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and does not equate to Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMainCode ExecutionMaterial ConstraintIntegrity
#1Doubao Pro96.794100pass
#2Gemini 2.5 Pro96.19795pass
#3GPT-o396.19795pass
#4Grok 496.19795pass
#5Qwen3 Max96.19795pass
#6Claude Opus 4.794.459495pass
#7Claude Sonnet 4.694.459495pass
#8DeepSeek V4 Pro94.459495warn
#9Gemini 3.1 Pro94.459495pass
#10GPT-5.594.459495pass

Data Interpretation

Today's Smoke quick test shows Doubao Pro leading the main ranking with 96.7, pairing 94 in code execution with a perfect 100 in material constraint. Gemini 2.5 Pro, GPT-o3, Grok 4, and Qwen3 Max all scored 96.1 on the main ranking, with a balanced structure of 97 in code execution and 95 in material constraint. Models from Claude Opus 4.7 to GPT-5.5 scored 94.45 on the main ranking, with 94 in code execution and 95 in material constraint, reflecting score distributions across different strength combinations.

Claude Sonnet 4.6 improved 39.4 points on the main ranking versus the previous run under the same methodology, with code execution up 42.4 points and material constraint up 35.8 points. Gemini 2.5 Pro improved 35.3 points on the main ranking, with code execution up 40.2 points and material constraint up 29.3 points. GPT-5.5, Grok 4, and GPT-o3 also posted main ranking gains ranging from 16.8 to 28.5 points. These changes may stem from question sampling fluctuations and require confirmation in subsequent runs.

GLM-4.6 is excluded from this round's ranking due to incomplete data (missing integrity and communication dimensions, API failure/timeout) and has entered automatic re-run. The Smoke quick test is a small-sample single-day signal, so interpretations of the above fluctuations should remain measured.

Key Changes

  • Claude Sonnet 4.6: Main ranking +39.4, code execution +42.4, material constraint +35.8
  • Gemini 2.5 Pro: Main ranking +35.3, code execution +40.2, material constraint +29.3
  • GPT-5.5: Main ranking +28.5, code execution +35.7, material constraint +19.6
  • Grok 4: Main ranking +26.5, code execution +22, material constraint +32.1
  • GPT-o3: Main ranking +16.8, code execution +15.2, material constraint +18.8

Signals to Watch

  • GLM-4.6: Incomplete data (missing integrity and communication dimensions, API failure/timeout), has entered automatic re-run, not ranked this round

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring verification in subsequent runs.


Data source: YZ Index | Run #257 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!