Claude Opus 4.7 Leads with 96.99: 2026-09-20 Smoke Quick-Test Data Brief

The 2026-09-20 YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first that day at 96.99. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to Full weekly ranking conclusions.

This Smoke evaluation covered only two main ranking dimensions: Code Execution and Material Constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capability.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.796.9910093.3pass
#2Doubao Pro94.3599.388.3pass
#3Claude Sonnet 4.689.5210076.7pass
#4Qwen3 Max89.5210076.7pass
#5GPT-5.586.510070pass
#6DeepSeek V4 Pro75.777576.7pass
#7Gemini 3.1 Pro75.3874.376.7pass
#8Grok 475.3874.376.7pass
#9GPT-o372.757570pass
#10Gemini 2.5 Pro67.245088.3pass

Data Interpretation

In terms of score structure, Claude Opus 4.7 scored 100 in Code Execution and 93.3 in Material Constraints to take a main ranking score of 96.99, while Doubao Pro scored 99.3 in Code Execution and 88.3 in Material Constraints for 94.35, showing that the two form a relatively balanced combination across the two metrics. Claude Sonnet 4.6 and Qwen3 Max both scored 100 in Code Execution and 76.7 in Material Constraints, tied at 89.52; GPT-5.5 scored 100 in Code Execution and 70 in Material Constraints, placing at 86.5. Leading models generally show high Code Execution and upper-middle Material Constraints.

DeepSeek V4 Pro rose 20.8 points on the main ranking compared with the previous run, with Code Execution up 50 points to 75 and Material Constraints down 15 points to 76.7; Doubao Pro rose 17.5 points on the main ranking and 34.8 points in Material Constraints to 88.3; Gemini 3.1 Pro rose 16.5 points on the main ranking and 24.3 points in Code Execution to 74.3. Grok 4 fell 13 points on the main ranking and 22.4 points in Code Execution to 74.3; Gemini 2.5 Pro fell 23.7 points in Code Execution to 50. GLM-4.6 did not participate in the ranking due to incomplete dimension data.

DeepSeek V4 Pro, whose Material Constraints plunged 15 points; Grok 4, whose main ranking plunged 13 points; and Gemini 2.5 Pro, whose Code Execution plunged 23.7 points, all show anomalous signals that may stem from single-day question sampling fluctuation or may indicate genuine performance degradation, requiring follow-up runs for confirmation. The Smoke quick test is a small-sample, single-day signal; the above interpretation remains restrained and makes no long-term judgment.

Key Changes

  • DeepSeek V4 Pro: main ranking up 20.8 points, Code Execution +50 points, Material Constraints -15 points
  • Doubao Pro: main ranking up 17.5 points, Material Constraints +34.8 points
  • Gemini 3.1 Pro: main ranking up 16.5 points, Code Execution +24.3 points, Material Constraints +6.9 points
  • Grok 4: main ranking down 13 points, Code Execution -22.4 points, integrity warn→pass
  • Qwen3 Max: main ranking up 12.5 points, Material Constraints +26.2 points, integrity warn→pass

Signals to Watch

  • DeepSeek V4 Pro: Material Constraints plunged by 15 points
  • Grok 4: main ranking plunged by 13 points
  • Gemini 2.5 Pro: Code Execution plunged by 23.7 points
  • GLM-4.6: incomplete data (missing execution, source attribution, judgment, integrity, and communication dimensions; API failure/timeout), has entered automatic backfill, and is not ranked this period

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be an early signal of genuine degradation, and require follow-up runs for review.


Data source: YZ Index | Run #330 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!