DeepSeek V4 Pro Ranks First with 87.2 Points: 2026-08-10 Smoke Quick Test Data Briefing

On 2026-08-10, the YZ Index Smoke quick test covered 10 models, and DeepSeek V4 Pro ranked first for the day with a score of 87.2. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the conclusions of the Full weekly ranking.

This Smoke evaluation only covers two main index dimensions: code execution and material constraint. The main index formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are more suitable for use as monitoring signals rather than as a basis for long-term conclusions about model capabilities.

Daily Rankings

RankModelMain IndexCode ExecutionMaterial ConstraintIntegrity
#1DeepSeek V4 Pro87.291.781.7pass
#2Grok 473.527571.7pass
#3Gemini 2.5 Pro72.8673.871.7pass
#4Claude Opus 4.772.37569pass
#5GPT-o371.277566.7pass
#6Gemini 3.1 Pro66.766.766.7pass
#7Qwen3 Max65.017552.8pass
#8GLM-4.661.255075pass
#9GPT-5.558.317537.9pass
#10Claude Sonnet 4.655.275061.7pass

Data Interpretation

In today's YZ Index Smoke quick test, DeepSeek V4 Pro ranked first with a main index score of 87.2, with code execution at 91.7 and material constraint at 81.7 forming a balanced high-level pairing, delivering outstanding overall performance under the 0.55 and 0.45 weights. Grok 4 (main index 73.52, code execution 75, material constraint 71.7), Gemini 2.5 Pro (main index 72.86, code execution 73.8, material constraint 71.7), Claude Opus 4.7 (main index 72.3, code execution 75, material constraint 69), and GPT-o3 (main index 71.27, code execution 75, material constraint 66.7) all exhibit a structural pattern in which code execution is slightly stronger than material constraint. GLM-4.6, by contrast, forms a reverse pairing with code execution at 50 and material constraint at 75, for a main index of 61.25; Qwen3 Max has code execution at 75 and material constraint at 52.8, for a main index of 65.01. Such complementary strengths and weaknesses are fairly common on the leaderboard.

Compared with the previous run under the same methodology, GPT-5.5's main index fell 30.4 points, code execution fell 25 points, and material constraint fell 37.1 points; GPT-o3's main index fell 24.6 points, code execution fell 25 points, and material constraint fell 24.2 points; Qwen3 Max's main index fell 20.8 points, code execution fell 24 points, and material constraint fell 16.9 points; Claude Opus 4.7's main index fell 20.6 points, code execution fell 25 points, and material constraint fell 15.1 points. GLM-4.6's main index rose 23.5 points, with material constraint up 52.2 points. These changes may stem from question sampling fluctuations, or may be incidental single-day signals, and require confirmation in subsequent runs. As a small-sample single-day signal, the current data is for reference only and does not constitute a basis for long-term judgment.

Key Changes

  • GPT-5.5: Main index -30.4 points, code execution -25 points, material constraint -37.1 points
  • GPT-o3: Main index -24.6 points, code execution -25 points, material constraint -24.2 points, integrity warn→pass
  • GLM-4.6: Main index +23.5 points, material constraint +52.2 points, integrity warn→pass
  • Qwen3 Max: Main index -20.8 points, code execution -24 points, material constraint -16.9 points
  • Claude Opus 4.7: Main index -20.6 points, code execution -25 points, material constraint -15.1 points

Signals to Watch

  • No publishable abnormal signals were retained this time.

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a particular model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling, or may be early signals of genuine degradation, requiring confirmation in subsequent runs.


Data source: YZ Index | Run #272 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!