Claude Opus 4.7 Tops with 81.57 Points: 2026-10-04 Smoke Quick-Test Data Brief

The 2026-10-04 YZ Index Smoke quick test covered 15 models, with Claude Opus 4.7 leading the day at 81.57 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, a single-day score is better suited as a monitoring signal rather than a long-term conclusion about model capability.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.781.579762.7pass
#2GPT-o379.4497.857pass
#3Doubao Pro75.579650.6pass
#4Grok 474.295.348.4pass
#5DeepSeek V4 Pro73.527571.7pass
#6Gemini 3.1 Pro72.537569.5pass
#7Qwen3 Max66.6961.672.9pass
#8GPT-6 Luna64.2769.857.5pass
#9Gemini 2.5 Pro63.6869.856.2pass
#10GPT-6.1 Sol62.5169.853.6pass
#11Claude Sonnet 4.662.3271.950.6pass
#12GPT-6 Sol61.1669.850.6pass
#13GPT-5.560.117541.9pass
#14GPT-6 Astra52.8644.862.7pass
#15GLM-4.633.945014.3warn

Data Interpretation

Among today's top three on the main leaderboard, Claude Opus 4.7 scored 81.57 with a combination of Code Execution 97 and Material Constraints 62.7; GPT-o3 reached 79.44 by relying on Code Execution 97.8 and Material Constraints 57; Doubao Pro ranked third with Code Execution 96 and Material Constraints 50.6. These models remain high on the Code Execution dimension while scoring relatively low on Material Constraints, creating a clearly imbalanced profile. DeepSeek V4 Pro and Gemini 3.1 Pro show more balanced combinations: the former has Code Execution 75 and Material Constraints 71.7, while the latter has Code Execution 75 and Material Constraints 69.5, with main leaderboard scores of 73.52 and 72.53, respectively.

Several models showed notable declines versus the previous same-basis run: GPT-6 Sol fell 24.1 points on the main leaderboard and 47.2 points on Material Constraints; Gemini 2.5 Pro fell 23.7 points on the main leaderboard and 30.2 points on Code Execution; GPT-6 Astra fell 21.9 points on the main leaderboard and 32.1 points on Material Constraints; Material Constraints for Claude Sonnet 4.6 and Claude Opus 4.7 fell 36.2 and 34.3 points, respectively. Smoke tests are small-sample, single-day signals; such changes may stem from question sampling fluctuations or may reflect genuine capability degradation, and require follow-up runs to confirm.

Overall, leading models mostly rely on Code Execution advantages to lift their main leaderboard scores, while combinations with weaker Material Constraints are relatively concentrated in today's data. Qwen3 Max's inverse structure—Code Execution 61.6 and Material Constraints 72.9—gives it a main leaderboard score of 66.69, showing the direct impact of different weighted combinations on ranking. All interpretations are based on exact values for the day and no extrapolation has been made.

Key Changes

  • GPT-6 Sol: Main leaderboard down 24.1 points, Code Execution -5.2 points, Material Constraints -47.2 points
  • Gemini 2.5 Pro: Main leaderboard down 23.7 points, Code Execution -30.2 points, Material Constraints -15.8 points
  • GPT-6 Astra: Main leaderboard down 21.9 points, Code Execution -13.5 points, Material Constraints -32.1 points
  • Claude Sonnet 4.6: Main leaderboard down 17.3 points, Material Constraints -36.2 points
  • Claude Opus 4.7: Main leaderboard down 17.1 points, Material Constraints -34.3 points

Signals to Watch

  • No publishable anomaly signals were retained for this run.

When reading this type of Smoke brief, focus on two questions: first, whether a model exposes the same type of weakness on multiple consecutive days; second, whether its Integrity rating moves from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.


Data source: YZ Index | Run #359 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!