DeepSeek V4 Pro, Gemini 3.1 Pro, and GLM-4.6 Tie at 83.49: 2026-09-07 Smoke Quick-Test Data Brief

On 2026-09-07, the YZ Index Smoke quick test covered 11 models, with DeepSeek V4 Pro, Gemini 3.1 Pro, and GLM-4.6 tying for the top spot of the day at 83.49 points. Smoke is a daily 10-question quick test suited to observing short-term signals and is not equivalent to the conclusions of the Full weekly leaderboard.

This Smoke evaluation only covers two main leaderboard dimensions: Code Execution and Material Constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term verdicts on model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintIntegrity
#1DeepSeek V4 Pro83.4910063.3pass
#2Gemini 3.1 Pro83.4910063.3pass
#3GLM-4.683.4910063.3warn
#4Doubao Pro80.8599.358.3pass
#5Grok 480.1399.356.7pass
#6GPT-o377.510050pass
#7GPT-5.572.037568.4pass
#8Gemini 2.5 Pro71.1673.568.3pass
#9Claude Sonnet 4.665.247553.3pass
#10Claude Opus 4.761.255075pass
#11Qwen3 Max55.995063.3pass

Key Changes

  • GLM-4.6: Main leaderboard up 19.8 points, code execution +30 points, material constraint +7.4 points, integrity pass→warn
  • Gemini 2.5 Pro: Main leaderboard down 17.1 points, code execution -16.6 points, material constraint -17.6 points
  • Doubao Pro: Main leaderboard up 14.7 points, code execution +49.3 points, material constraint -27.6 points
  • Grok 4: Main leaderboard up 11.5 points, code execution +24.3 points
  • DeepSeek V4 Pro: Main leaderboard up 6.1 points, code execution +8.3 points

Signals to Watch

  • Doubao Pro: Material constraint plunged -27.6 points
  • GPT-o3: Material constraint plunged -20 points
  • Gemini 2.5 Pro: Main leaderboard plunged -17.1 points

When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness across multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signs of genuine degradation that require follow-up review in subsequent runs.


Data source: YZ Index | Run #312 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!