Claude Opus 4.7 Leads with 100 Points: 2026-07-20 Smoke Quick Test Data Brief

Claude Opus 4.7 Leads with 100 Points: 2026-07-20 Smoke Quick Test Data Brief

On 2026-07-20, the YZ Index Smoke Quick Test covered 11 models, with Claude Opus 4.7 scoring 100 points to top the daily rankings. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.7100100100pass
#2Doubao Pro83.737594.4pass
#3GPT-o382.77592.1pass
#4Gemini 3.1 Pro81.937590.4pass
#5Grok 480.817587.9warn
#6DeepSeek V4 Pro80.277586.7warn
#7GPT-5.579.917585.9pass
#8Gemini 2.5 Pro69.985094.4pass
#9Qwen3 Max67.3165.669.4warn
#10Claude Sonnet 4.664.3971.955.2pass
#11GLM-4.658.7525100warn

Data Interpretation

Among today's top four models in the main ranking, Claude Opus 4.7 achieved a main score of 100 with both code execution and material constraints at 100, showing a balanced high level. Doubao Pro's main score of 83.73 came from code execution at 75 and material constraints at 94.4; GPT-o3's main score of 82.7 corresponded to code execution at 75 and material constraints at 92.1; Gemini 3.1 Pro's main score of 81.93 came from code execution at 75 and material constraints at 90.4. This structure indicates that combinations with relatively stronger material constraints contributed higher overall scores in the current sample.

Doubao Pro's main score increased by 34.5 points compared to the previous run under the same conditions, mainly driven by a 50-point rise in code execution. GLM-4.6's main score rose 27.3 points, accompanied by a 60.7-point increase in material constraints. Gemini 3.1 Pro's main score increased by 25.5 points, with code execution up 25 points and material constraints up 26.1 points. Regarding anomalies, Gemini 2.5 Pro's code execution saw a sharp drop of -24.6 points, Qwen3 Max's main score plunged -14.9 points, and Claude Sonnet 4.6's main score plummeted -25.6 points. These single-day changes may stem from question sampling fluctuations or could indicate real performance degradation, requiring confirmation from subsequent runs under the same conditions.

The Smoke test is a small-sample single-day signal; the above observations only reflect the day's data distribution and do not make judgments on long-term model stability.

Key Changes

  • Doubao Pro: Main score +34.5, Code Execution +50, Material Constraints +15.6
  • GLM-4.6: Main score +27.3, Material Constraints +60.7
  • Claude Sonnet 4.6: Main score -25.6, Code Execution -27.3, Material Constraints -23.6
  • Gemini 3.1 Pro: Main score +25.5, Code Execution +25, Material Constraints +26.1
  • DeepSeek V4 Pro: Main score +17.3, Code Execution +25, Material Constraints +7.9, Integrity pass→warn

Signals to Watch

  • Gemini 2.5 Pro: Code Execution plunged -24.6 points
  • Qwen3 Max: Main score plunged -14.9 points
  • Claude Sonnet 4.6: Main score plunged -25.6 points

When reading such Smoke briefs, the focus should be on two questions: first, whether a particular model exposes the same type of weakness on consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may stem from question sampling or could be early signals of real degradation, requiring confirmation from subsequent runs.


Data source: YZ Index | Run #238 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!