Claude Opus 4.7, DeepSeek V4 Pro, Doubao Pro, and GPT-5.5 Tie at 91.09: 2026-09-14 Smoke Quick Test Data Brief

The 2026-09-14 YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7, DeepSeek V4 Pro, Doubao Pro, and GPT-5.5 tying for first place that day at 91.09. Smoke is a daily 10-question quick test, suited to observing short-term signals; it is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Claude Opus 4.791.0910080.2pass
#2DeepSeek V4 Pro91.0910080.2pass
#3Doubao Pro91.0910080.2pass
#4GPT-5.591.0910080.2pass
#5Gemini 3.1 Pro8710071.1pass
#6Grok 486.6199.371.1pass
#7Claude Sonnet 4.682.510061.1pass
#8Qwen3 Max70.4762.580.2warn
#9Gemini 2.5 Pro70.197564.3pass
#10GPT-o370.197564.3pass

Data Interpretation

The top four models on today's main leaderboard—Claude Opus 4.7, DeepSeek V4 Pro, Doubao Pro, and GPT-5.5—all show a structure of 100 in code execution and 80.2 in material constraints, with the same main leaderboard score of 91.09, indicating a balanced mix of strengths and weaknesses with consistent values. Gemini 3.1 Pro also scores 100 in code execution but drops to 71.1 in material constraints, with its main leaderboard score falling to 87; Grok 4 scores 99.3 in code execution and 71.1 in material constraints, for a main leaderboard score of 86.61; Claude Sonnet 4.6 scores 100 in code execution and 61.1 in material constraints, for a main leaderboard score of 82.5, reflecting the impact of the material constraints gap on overall placement.

GPT-5.5's main leaderboard score rose 30.2 points from the previous same-scope run, with code execution up 25 points and material constraints up 36.6 points; Claude Sonnet 4.6's main leaderboard score rose 28.9 points, with code execution up 50 points; Claude Opus 4.7's main leaderboard score rose 27.5 points, with code execution up 50 points; DeepSeek V4 Pro's main leaderboard score rose 13.8 points, with code execution up 25 points; GPT-o3's main leaderboard score fell 15 points, with code execution down 25 points. These changes all come from single-day Smoke small-sample data and require follow-up runs to distinguish question-sampling fluctuations from genuine degradation.

GPT-o3 shows an anomalous -15-point signal on the main leaderboard, which may stem from single-day sampling fluctuation or a temporary change in model performance and requires a full run for verification; GLM-4.6 did not participate in the ranking because its data was incomplete (missing several evaluation dimensions; API failure or timeout). An automatic rerun has been triggered, and this period's results are for daily reference only.

Key Changes

  • GPT-5.5: main leaderboard up 30.2 points, code execution +25 points, material constraints +36.6 points
  • Claude Sonnet 4.6: main leaderboard up 28.9 points, code execution +50 points
  • Claude Opus 4.7: main leaderboard up 27.5 points, code execution +50 points
  • GPT-o3: main leaderboard down 15 points, code execution -25 points
  • DeepSeek V4 Pro: main leaderboard up 13.8 points, code execution +25 points

Signals to Watch

  • GPT-o3: main leaderboard showed a sharp -15-point drop
  • GLM-4.6: incomplete data (missing several evaluation dimensions; API failure/timeout), has entered automatic rerun, and does not participate in this period's ranking

When reading this kind of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs for verification.


Data Source: YZ Index | Run #322 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!