Gemini 3.1 Pro Tops with 96.98 Points: 2026-08-24 Smoke Quick Test Data Briefing

The YZ Index Smoke quick test on 2026-08-24 covered 11 models, with Gemini 3.1 Pro ranking first at 96.98 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintIntegrity
#1Gemini 3.1 Pro96.9894.5100pass
#2Gemini 2.5 Pro92.1694.589.3pass
#3GLM-4.687.5486.189.3warn
#4Grok 487.0986.188.3pass
#5Doubao Pro86.5494.576.8warn
#6GPT-o383.2369.5100pass
#7Claude Opus 4.782.9877.889.3pass
#8DeepSeek V4 Pro80.4694.563.3pass
#9Qwen3 Max71.7377.864.3pass
#10GPT-5.564.6644.589.3pass
#11Claude Sonnet 4.658.2344.575pass

Data Interpretation

In today's YZ Index Smoke quick test, Gemini 3.1 Pro achieved a main ranking of 96.98 with a combination of 94.5 in code execution and 100 in material constraint. Gemini 2.5 Pro's scores of 94.5 in code execution and 89.3 in material constraint corresponded to a main ranking of 92.16. GLM-4.6 scored 87.54 on the main ranking with 86.1 in code execution and 89.3 in material constraint, showing clear differences in the strength combinations of top models across code execution and material constraint. Doubao Pro scored 94.5 in code execution but 76.8 in material constraint, with a main ranking of 86.54; GPT-o3 scored 69.5 in code execution and 100 in material constraint, with a main ranking of 83.23. This type of structure features strength at one end and limitation at the other.

Compared with the previous run using the same methodology, GLM-4.6 gained +64.1 points on the main ranking, +69.4 on code execution, and +57.6 on material constraint; Gemini 2.5 Pro gained +28.9 on the main ranking, +25 on code execution, and +33.6 on material constraint; GPT-o3 gained +27.7 on the main ranking, +19.5 on code execution, and +37.6 on material constraint; Grok 4 gained +22.8 on the main ranking, +19.4 on code execution, and +26.9 on material constraint; Gemini 3.1 Pro gained +18.5 on the main ranking and +37.6 on material constraint. These increases reflect changes in the single-day score structure.

DeepSeek V4 Pro's material constraint score plummeted by -18.4 points, which is an anomalous signal that may stem from question sampling fluctuations or genuine degradation, requiring confirmation through subsequent run reviews. The Smoke quick test provides small-sample, single-day signals. The above interpretation is based solely on today's data and makes no long-term judgments.

Key Changes

  • GLM-4.6: main ranking up 64.1 points, code execution +69.4, material constraint +57.6
  • Gemini 2.5 Pro: main ranking up 28.9 points, code execution +25, material constraint +33.6
  • GPT-o3: main ranking up 27.7 points, code execution +19.5, material constraint +37.6
  • Grok 4: main ranking up 22.8 points, code execution +19.4, material constraint +26.9
  • Gemini 3.1 Pro: main ranking up 18.5 points, material constraint +37.6

Signals to Watch

  • DeepSeek V4 Pro: material constraint plummeted -18.4 points

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Significant single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring review in subsequent runs.


Data Source: YZ Index | Run #292 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!