Gemini 3.1 Pro Leads at 94.74: 2026-08-11 Smoke Quick Test Data Briefing

On 2026-08-11, the YZ Index Smoke quick test covered 10 models, with Gemini 3.1 Pro ranking first that day at 94.74 points. Smoke is a daily 10-question quick test suited for observing short-term signals; it is not equivalent to the Full weekly ranking conclusion.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × code execution + 0.45 × material constraints. Given the small daily sample size, single-day scores are better used as monitoring signals rather than as long-term verdicts on model capabilities.

Daily Ranking

RankModelMain ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Gemini 3.1 Pro94.7410088.3pass
#2Claude Opus 4.792.9410084.3pass
#3DeepSeek V4 Pro87.9910073.3pass
#4Grok 487.9395.878.3pass
#5Gemini 2.5 Pro86.2892.878.3pass
#6Claude Sonnet 4.685.7410068.3pass
#7GPT-o384.4194.871.7pass
#8GPT-5.583.1710062.6pass
#9Qwen3 Max81.5687.574.3pass
#10GLM-4.6747578.3fail

Key Changes

  • Claude Sonnet 4.6: Main Score up 30.5 points, Code Execution +50 points, Material Constraints +6.6 points
  • Gemini 3.1 Pro: Main Score up 28 points, Code Execution +33.3 points, Material Constraints +21.6 points
  • GPT-5.5: Main Score up 24.9 points, Code Execution +25 points, Material Constraints +24.7 points
  • Claude Opus 4.7: Main Score up 20.6 points, Code Execution +25 points, Material Constraints +15.3 points
  • Qwen3 Max: Main Score up 16.6 points, Code Execution +12.5 points, Material Constraints +21.5 points

Signals to Watch

  • GLM-4.6: Integrity rating downgraded to Fail (pass→fail)
  • Doubao Pro: Incomplete data (missing execution, material-constraint, judgment, integrity, and communication dimensions; API failure/timeout). Automatic re-run initiated; not ranked in this round.

When reading Smoke briefings of this kind, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling, or they may be early signals of genuine regression, requiring review in subsequent runs.


Data source: YZ Index | Run #274 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!