Claude Opus 4.7 Leads with 100 Points: 2026-09-18 Smoke Quick Test Data Brief

On 2026-09-18, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 ranking first for the day with 100 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals and not equivalent to Full weekly ranking conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are more suitable as monitoring signals than as long-term conclusions about model capabilities.

Today's Ranking

RankModelMainCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.7100100100pass
#2GPT-o399.0110097.8pass
#3Claude Sonnet 4.694.3499.687.9pass
#4GPT-5.592.269588.9pass
#5Gemini 2.5 Pro90.9110079.8pass
#6Doubao Pro83.570100warn
#7Qwen3 Max82.4289.374pass
#8Grok 481.177588.7pass
#9Gemini 3.1 Pro69.7545100pass
#10DeepSeek V4 Pro67.515088.9pass

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first with a structure of Main 100, Code Execution 100, and Material Constraints 100, forming a completely balanced combination of Code Execution and Material Constraints. GPT-o3 had Main 99.01, Code Execution 100, and Material Constraints 97.8, showing that Code Execution remained at a full score while Material Constraints were slightly lower. Claude Sonnet 4.6 had Main 94.34, Code Execution 99.6, and Material Constraints 87.9, likewise showing a combination in which Code Execution is stronger than Material Constraints. Gemini 2.5 Pro had Main 90.91, Code Execution 100, and Material Constraints 79.8, further highlighting the characteristic of high Code Execution and low Material Constraints. By contrast, Doubao Pro had Main 83.5, Code Execution 70, and Material Constraints 100, reflecting a complementary structure with a full score in Material Constraints but relatively weak Code Execution.

In terms of significant changes, Claude Opus 4.7's Main score rose 33.2 points compared with the previous same-basis run, with Code Execution up 53 points and Material Constraints up 9.1 points; Qwen3 Max's Main score rose 27.8 points, with Code Execution up 64.3 points but Material Constraints down 16.9 points; GPT-o3's Main score rose 21.6 points, with Code Execution up 28 points and Material Constraints up 13.7 points. These increases were mainly driven by the Code Execution dimension, while Material Constraints diverged. Qwen3 Max's 16.9-point plunge in Material Constraints may stem from single-day question sampling fluctuation, and may also represent real degradation; subsequent runs are needed to verify signal stability. GLM-4.6 did not participate in this period's ranking because an API failure led to incomplete data.

Overall, leading models have different emphases in the strength mix between Code Execution and Material Constraints; score changes among the models with notable moves are concentrated in the Code Execution dimension, while fluctuations on the Material Constraints side are relatively limited. As a small-sample single-day signal, Smoke's observations above only reflect that day's data characteristics and do not constitute a basis for long-term judgment.

Main Changes

  • Claude Opus 4.7: Main up 33.2 points, Code Execution +53 points, Material Constraints +9.1 points
  • Qwen3 Max: Main up 27.8 points, Code Execution +64.3 points, Material Constraints -16.9 points
  • GPT-o3: Main up 21.6 points, Code Execution +28 points, Material Constraints +13.7 points
  • DeepSeek V4 Pro: Main up 19.7 points, Material Constraints +43.7 points
  • Claude Sonnet 4.6: Main up 17.9 points, Code Execution +27.6 points, Material Constraints +6.1 points

Signals to Watch

  • Qwen3 Max: Material Constraints plunge of -16.9 points
  • GLM-4.6: Incomplete data (missing execution, judgment, integrity, and communication dimensions; API failure/timeout); an automatic rerun has been initiated, and it does not participate in this period's ranking

When reading this kind of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be an early signal of real degradation, requiring subsequent runs for review.


Data source: YZ Index | Run #328 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!