DeepSeek V4 Pro, GPT-5.5, and GPT-o3 Tie at 80.52 Points: 2026-08-05 Smoke Quick Test Data Briefing

On 2026-08-05, the YZ Index Smoke quick test covered 9 models, with DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tying for first place at 80.52 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusion.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintsIntegrity
#1DeepSeek V4 Pro80.5210056.7pass
#2GPT-5.580.5210056.7pass
#3GPT-o380.5210056.7pass
#4Claude Sonnet 4.674.0996.946.2pass
#5Gemini 2.5 Pro73.7481.364.5pass
#6Claude Opus 4.771.947568.2pass
#7Gemini 3.1 Pro55.527531.7pass
#8Grok 455.527531.7pass
#9Qwen3 Max49.7271.922.6pass

Data Interpretation

In today's YZ Index Smoke quick test, DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tied on the main ranking at 80.52, each scoring 100 in code execution and 56.7 in material constraints, showing a structural combination of perfect code execution with moderate material constraints. Claude Sonnet 4.6 scored 74.09 on the main ranking, with 96.9 in code execution and 46.2 in material constraints, also exhibiting strong code execution and weaker material constraints. Gemini 2.5 Pro scored 73.74 on the main ranking, with 81.3 in code execution and 64.5 in material constraints, forming a relatively balanced but overall lower combination across both metrics. Claude Opus 4.7 scored 71.94 on the main ranking, with 75 in code execution and 68.2 in material constraints, where material constraints exceed code execution, forming another type of combination.

Compared with the previous run under the same methodology, Qwen3 Max fell 35.1 points on the main ranking, with code execution down 20.8 points and material constraints down 52.6 points; Grok 4 fell 28.8 points on the main ranking, with code execution down 25 points and material constraints down 33.5 points; Gemini 3.1 Pro fell 27.6 points on the main ranking, with code execution down 22 points and material constraints down 34.4 points. These declines are concentrated in material constraints and code execution, and subsequent runs are needed to distinguish between question-sampling fluctuations and genuine regression. The Smoke quick test is a small-sample single-day signal; the current data only reflects that day's performance and does not constitute a basis for long-term judgment.

Overall, top models largely rely on high code execution scores to support their main ranking, while material constraint scores are generally lower, creating a clear structural disparity. The synchronized decline among models showing anomalies suggests possible single-day test fluctuation, requiring multiple rounds of retesting to confirm stability.

Key Changes

  • Qwen3 Max: Main ranking down 35.1 points; code execution -20.8; material constraints -52.6
  • Grok 4: Main ranking down 28.8 points; code execution -25; material constraints -33.5
  • Gemini 3.1 Pro: Main ranking down 27.6 points; code execution -22; material constraints -34.4
  • Claude Opus 4.7: Main ranking down 17.5 points; code execution -22; material constraints -12
  • Gemini 2.5 Pro: Main ranking down 15.8 points; code execution -18.7; material constraints -12.3

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large daily swings in execution or constraint scores may stem from question sampling or may be early signals of genuine regression, requiring subsequent run verification.


Data source: YZ Index | Run #262 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!