Claude Opus 4.7 and GPT-o3 Tie at 100 Points: 2026-08-27 Smoke Quick Test Data Briefing

On 2026-08-27, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and GPT-o3 tied for first place at 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions: code execution and material constraint. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraint. Given the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.7100100100pass
#2GPT-o3100100100pass
#3GPT-5.595.0110088.9pass
#4Grok 493.4696.789.5pass
#5Claude Sonnet 4.690.7810079.5pass
#6Doubao Pro90.7810079.5pass
#7Gemini 3.1 Pro88.7510075pass
#8Gemini 2.5 Pro87.5999.273.4pass
#9GLM-4.686.2575100pass
#10Qwen3 Max84.7492.775pass
#11DeepSeek V4 Pro49.032578.4pass

Key Changes

  • GLM-4.6: Main leaderboard up 38.6 points, Code Execution up 37.5, Material Constraint up 40
  • DeepSeek V4 Pro: Main leaderboard down 35 points, Code Execution down 50, Material Constraint down 16.6
  • Claude Opus 4.7: Main leaderboard up 20.1 points, Code Execution up 25, Material Constraint up 14.1
  • GPT-5.5: Main leaderboard up 11 points, Code Execution up 25, Material Constraint down 6.1
  • Gemini 2.5 Pro: Main leaderboard down 8.5 points, Material Constraint down 21.6

Signals to Watch

  • Claude Sonnet 4.6: Material Constraint plummeted 20.5 points
  • Doubao Pro: Material Constraint plummeted 20.5 points
  • Gemini 3.1 Pro: Material Constraint plummeted 20 points
  • Gemini 2.5 Pro: Main leaderboard plummeted 8.5 points
  • DeepSeek V4 Pro: Main leaderboard plummeted 35 points

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large swings in single-day execution or constraint scores may stem from question sampling, or may be early signals of genuine degradation, requiring follow-up runs for verification.


Data source: YZ Index | Run #297 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!