GPT-o3 Tops with 91.29 Points: 2026-07-27 Smoke Quick Test Data Brief

On 2026-07-27, the YZ Index Smoke quick test covered 11 models, with GPT-o3 ranking first on the day with a score of 91.29. Smoke is a daily quick test of 10 questions, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers the two main ranking dimensions: Code Execution and Material Constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain RankingCode ExecutionMaterial ConstraintsIntegrity
#1GPT-o391.299784.3pass
#2Gemini 3.1 Pro89.049779.3pass
#3GPT-5.586.299773.2pass
#4DeepSeek V4 Pro85.6910068.2pass
#5Doubao Pro84.949770.2pass
#6Gemini 2.5 Pro84.2895.870.2pass
#7Claude Sonnet 4.680.449760.2pass
#8GLM-4.675.697280.2warn
#9Grok 475.1579.270.2pass
#10Claude Opus 4.766.044789.3pass
#11Qwen3 Max62.0759.565.2pass

Data Interpretation

Today’s YZ Index Smoke quick test shows that top models have different emphases in the combination of Code Execution and Material Constraints. GPT-o3 has a main ranking of 91.29, with Code Execution 97 and Material Constraints 84.3, showing a relatively balanced capability; Gemini 3.1 Pro has a main ranking of 89.04, Code Execution 97, Material Constraints 79.3, also relying on high Code Execution scores; DeepSeek V4 Pro has Code Execution 100, Material Constraints 68.2, main ranking 85.69, reflecting a structure with clear Code Execution advantage but weaker Material Constraints; GPT-5.5 and Doubao Pro both have Code Execution 97, Material Constraints 73.2 and 70.2 respectively, overall ranking high on the main ranking.

Among significant changes, GPT-5.5 main ranking +26.5 points, Code Execution +52.5 points; GPT-o3 main ranking +21.8 points, Code Execution +52.5 points; Qwen3 Max main ranking +18.9 points, Code Execution +40 points; Claude Sonnet 4.6 main ranking +15.7 points, Code Execution +52.5 points. These improvements were accompanied by varying degrees of decline in Material Constraints. Abnormal signals include GPT-o3 Material Constraints plummeting -15.7 points, DeepSeek V4 Pro Material Constraints plummeting -31.8 points, Doubao Pro Material Constraints plummeting -17 points, Claude Sonnet 4.6 Material Constraints plummeting -29.3 points, Grok 4 Material Constraints plummeting -29.8 points, and GLM-4.6 Integrity rating downgraded to warn, which may be due to question sampling fluctuations or genuine degradation, requiring confirmation in subsequent runs. Smoke is a small-sample single-day signal, and interpretations should remain restrained.

Major Changes

  • GPT-5.5: Main ranking +26.5 points, Code Execution +52.5 points, Material Constraints -5.2 points
  • GPT-o3: Main ranking +21.8 points, Code Execution +52.5 points, Material Constraints -15.7 points
  • Qwen3 Max: Main ranking +18.9 points, Code Execution +40 points, Material Constraints -6.8 points
  • GLM-4.6: Main ranking +18.5 points, Code Execution +27.5 points, Material Constraints +7.4 points, Integrity fail→warn
  • Claude Sonnet 4.6: Main ranking +15.7 points, Code Execution +52.5 points, Material Constraints -29.3 points

Signals to Watch

  • GPT-o3: Material Constraints plummeted -15.7 points
  • DeepSeek V4 Pro: Material Constraints plummeted -31.8 points
  • Doubao Pro: Material Constraints plummeted -17 points
  • Claude Sonnet 4.6: Material Constraints plummeted -29.3 points
  • GLM-4.6: Integrity rating downgraded to Fail (fail→warn)
  • Grok 4: Material Constraints plummeted -29.8 points

When reading this type of Smoke brief, the focus should be on two questions: first, whether a certain model has exposed the same type of weakness for consecutive days; second, whether the Integrity rating has moved from pass to warn or fail. Large daily changes in Execution or Constraint scores may come from question sampling or be early signals of genuine degradation, requiring confirmation in subsequent runs.


Data source: YZ Index | Run #248 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!