DeepSeek V4 Pro Ranks First with 82.64 Points: 2026-08-23 Smoke Quick Test Data Briefing

The 2026-08-23 YZ Index Smoke quick test covered 11 models, with DeepSeek V4 Pro ranking first at 82.64 points. Smoke is a daily 10-question quick test suited for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × code execution + 0.45 × material constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain RankingCode ExecutionMaterial ConstraintsIntegrity
#1DeepSeek V4 Pro82.6483.481.7pass
#2Gemini 3.1 Pro78.5291.762.4pass
#3Claude Opus 4.768.8358.381.7pass
#4Doubao Pro68.3266.770.3pass
#5Grok 464.3266.761.4pass
#6Gemini 2.5 Pro63.2969.555.7pass
#7GPT-5.558.155068.1pass
#8GPT-o355.585062.4pass
#9Qwen3 Max54.0358.348.8pass
#10Claude Sonnet 4.651.4658.343.1pass
#11GLM-4.623.4516.731.7warn

Data Interpretation

DeepSeek V4 Pro ranks first with a main ranking score of 82.64, with a balanced profile of 83.4 in code execution and 81.7 in material constraints. Gemini 3.1 Pro scores 78.52 on the main ranking, with 91.7 in code execution versus 62.4 in material constraints, showing a clear emphasis on code execution. Claude Opus 4.7 scores 68.83 on the main ranking, with 58.3 in code execution and 81.7 in material constraints, making material constraints relatively prominent. Doubao Pro has a main ranking score of 68.32, with a mid-range combination of 66.7 in code execution and 70.3 in material constraints. Grok 4 and Gemini 2.5 Pro post main ranking scores of 64.32 and 63.29 respectively, with code execution higher than material constraints for both.

GLM-4.6's main ranking dropped 48.1 points, code execution dropped 33.3 points, material constraints dropped 66.1 points, and integrity went from pass to warn. Qwen3 Max's main ranking dropped 44.6 points, code execution dropped 41.7 points, and material constraints dropped 48.2 points. Claude Sonnet 4.6's main ranking dropped 33.4 points, with material constraints dropping 53.9 points. These changes may stem from question sampling fluctuations or genuine single-day performance differences, and require confirmation in subsequent run reviews.

Claude Opus 4.7's main ranking dropped 31.2 points, and Doubao Pro's main ranking dropped 30.7 points, both accompanied by declines in both code execution and material constraints. The Smoke quick test is a small-sample single-day signal; the current data only reflects that day's results and does not constitute a basis for long-term judgment.

Key Changes

  • GLM-4.6: Main ranking down 48.1 points, code execution -33.3 points, material constraints -66.1 points, integrity pass→warn
  • Qwen3 Max: Main ranking down 44.6 points, code execution -41.7 points, material constraints -48.2 points
  • Claude Sonnet 4.6: Main ranking down 33.4 points, code execution -16.7 points, material constraints -53.9 points
  • Claude Opus 4.7: Main ranking down 31.2 points, code execution -41.7 points, material constraints -18.3 points
  • Doubao Pro: Main ranking down 30.7 points, code execution -33.3 points, material constraints -27.5 points

Signals to Watch

  • No anomalous signals were retained for publication this time.

When reading Smoke briefings like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores could stem from question sampling or could be early signals of genuine degradation, requiring review in subsequent runs.


Data source: YZ Index | Run #290 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!