Claude Opus 4.7 and GPT-5.5 Tie at 86.5: 2026-07-30 Smoke Quick Test Data Brief

Claude Opus 4.7 and GPT-5.5 Tie at 86.5: 2026-07-30 Smoke Quick Test Data Brief

On 2026-07-30, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and GPT-5.5 tying for first at 86.5 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covered two main board dimensions: Code Execution and Material Constraint. The main board formula is 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelOverall ScoreCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.786.510070pass
#2GPT-5.586.510070pass
#3GPT-o379.917585.9pass
#4Claude Sonnet 4.678.8696.956.8pass
#5Grok 478.019260.9pass
#6DeepSeek V4 Pro76.857579.1pass
#7Doubao Pro757575pass
#8Gemini 3.1 Pro72.757570pass
#9Gemini 2.5 Pro60.7973.745pass
#10Qwen3 Max60.5554.767.7pass
#11GLM-4.639.565026.8pass

Key Changes

  • GPT-5.5: Overall Score up 19.8 points, Code Execution +25 points, Material Constraint +13.5 points
  • Grok 4: Overall Score down 11.3 points, Code Execution -5.8 points, Material Constraint -18 points
  • Claude Opus 4.7: Overall Score up 9.7 points, Code Execution +25 points, Material Constraint -8.9 points
  • Claude Sonnet 4.6: Overall Score up 8.1 points, Code Execution +21.9 points, Material Constraint -8.8 points
  • Qwen3 Max: Overall Score down 7.5 points, Code Execution -17.8 points, Material Constraint +5.1 points

Signals to Watch

  • Grok 4: Overall Score plummeted -11.3 points
  • DeepSeek V4 Pro: Code Execution plummeted -25 points
  • Doubao Pro: Code Execution plummeted -22 points
  • Gemini 2.5 Pro: Material Constraint plummeted -24.8 points
  • Qwen3 Max: Code Execution plummeted -17.8 points
  • GLM-4.6: Material Constraint plummeted -40.7 points

When reading such Smoke briefs, the focus should be on two questions: First, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Significant daily fluctuations in execution or constraint scores may be due to item sampling or could be early signals of real degradation, requiring subsequent run verification.


Data source: YZ Index | Run #254 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!