Claude Opus 4.7 Tops at 93.54 Points: 2026-08-28 Smoke Quick Test Data Brief

On 2026-08-28, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first that day at 93.54 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain RankingCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.793.549789.3pass
#2GPT-o389.9210077.6pass
#3Gemini 2.5 Pro88.5710074.6pass
#4Doubao Pro86.929774.6pass
#5Qwen3 Max86.8799.271.8pass
#6GPT-5.585.939772.4pass
#7Grok 485.4996.272.4warn
#8Claude Sonnet 4.683.779767.6pass
#9Gemini 3.1 Pro63.5554.574.6pass
#10GLM-4.661.075074.6pass
#11DeepSeek V4 Pro58.434772.4pass

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 ranked first with a main ranking score of 93.54, showing the most balanced combination of code execution at 97 and material constraint at 89.3. GPT-o3 and Gemini 2.5 Pro both stood out with perfect code execution scores of 100, but their material constraint scores of 77.6 and 74.6 respectively widened the gap in main ranking to 89.92 and 88.57. Doubao Pro and Qwen3 Max posted code execution scores of 97 and 99.2 and material constraint scores of 74.6 and 71.8, also exhibiting a structural profile stronger on the code side.

Gemini 3.1 Pro fell 25.2 points in the main ranking and 45.5 points in code execution; GLM-4.6 also fell 25.2 points in the main ranking, with code execution down 25 points and material constraint down 25.4 points; GPT-o3 fell 10.1 points in the main ranking and 22.4 points in material constraint. These anomalies may stem from sampling variance in a single day's questions or may reflect temporary performance changes of models in specific constraint scenarios, requiring confirmation through subsequent runs under the same methodology. Grok 4 and GPT-5.5 saw main ranking drops of 8 points and 9.1 points respectively, also indicating that volatility on the material constraint side warrants continued observation.

Overall, top-tier models maintained stable strength profiles across code execution and material constraint, while the significant score fluctuations among mid-to-lower-tier models serve as a reminder that the Smoke quick test is merely a small-sample, single-day signal, and any interpretation should be treated with caution.

Key Changes

  • Gemini 3.1 Pro: Main ranking down 25.2 points, code execution -45.5 points
  • GLM-4.6: Main ranking down 25.2 points, code execution -25 points, material constraint -25.4 points
  • GPT-o3: Main ranking down 10.1 points, material constraint -22.4 points
  • DeepSeek V4 Pro: Main ranking up 9.4 points, code execution +22 points, material constraint -6 points
  • GPT-5.5: Main ranking down 9.1 points, material constraint -16.5 points

Signals to Watch

  • GPT-o3: Main ranking plunged -10.1 points
  • GPT-5.5: Main ranking plunged -9.1 points
  • Grok 4: Main ranking plunged -8 points
  • Gemini 3.1 Pro: Main ranking plunged -25.2 points
  • GLM-4.6: Main ranking plunged -25.2 points

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large single-day swings in execution or constraint scores may stem from question sampling or could be early signs of genuine degradation, requiring verification in subsequent runs.


Data source: YZ Index | Run #298 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!