Claude Opus 4.7 and Qwen3 Max Tie at 93.39: 2026-08-01 Smoke Quick Test Data Brief

On 2026-08-01, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and Qwen3 Max tying for first place at 93.39 points. Smoke is a daily 10-question quick test suited for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main ranking dimensions: code execution and material constraint. The main ranking formula is 0.55 × code execution + 0.45 × material constraint. Due to the small daily sample size, single-day scores are better used as monitoring signals rather than as definitive long-term conclusions about model capability.

Daily Ranking

RankModelOverallCode ExecutionMaterial ConstraintIntegrity
#1Claude Opus 4.793.3910085.3pass
#2Qwen3 Max93.3910085.3pass
#3DeepSeek V4 Pro90.9998.581.8pass
#4Doubao Pro88.6210074.7pass
#5Gemini 3.1 Pro82.167590.9pass
#6GPT-o379.2881.876.2pass
#7GLM-4.676.4710047.7warn
#8Grok 469.567562.9pass
#9GPT-5.56658.375.4pass
#10Gemini 2.5 Pro60.8156.865.7pass
#11Claude Sonnet 4.655.0251.659.2pass

Data Interpretation

In today's YZ Index Smoke quick test, Claude Opus 4.7 and Qwen3 Max tied at 93.39 overall, both scoring 100 in code execution and 85.3 in material constraint, presenting a balanced high-level combination of code execution and material constraint. DeepSeek V4 Pro scored 90.99 overall, with code execution at 98.5 and material constraint at 81.8, also maintaining a strong code execution advantage. Doubao Pro scored 88.62 overall, with code execution at 100 but material constraint at 74.7, showing a structural profile of strong code execution paired with relatively weaker material constraint. Gemini 3.1 Pro scored 82.16 overall and showed the opposite pattern — code execution at 75 and material constraint at 90.9 — with material constraint serving as its main support.

GLM-4.6 scored 76.47 overall, with code execution at 100 but material constraint at only 47.7, and its integrity rating shifted from pass to warn. Its material constraint changed by -27.3 points compared with the previous run. Qwen3 Max rose 21.1 points overall and 37.6 points in material constraint versus the previous run. GPT-5.5 scored 66 overall, with code execution at 58.3 and material constraint at 75.4, down 20.6 points overall and 36.7 points in code execution from the previous run. Claude Sonnet 4.6 scored 55.02 overall, with code execution at 51.6 and material constraint at 59.2, down 19 points overall and 48.4 points in code execution from the previous run.

On anomaly signals: GPT-o3 fell 13.9 points overall and 14.7 points in material constraint; GLM-4.6 fell 27.3 points in material constraint; Grok 4 fell 8.2 points overall; GPT-5.5 fell 20.6 points overall; Gemini 2.5 Pro fell 8.5 points overall; and Claude Sonnet 4.6 fell 19 points overall. These single-day fluctuations may stem from question sampling variance or may reflect genuine performance changes, and require confirmation in subsequent runs under the same methodology. The Smoke quick test is a small-sample, single-day signal; the above observations are provided for reference for that day's data only.

Key Changes

  • GLM-4.6: Overall +30.2 points, code execution +77.2 points, material constraint -27.3 points, integrity pass→warn
  • Qwen3 Max: Overall +21.1 points, code execution +7.5 points, material constraint +37.6 points
  • GPT-5.5: Overall -20.6 points, code execution -36.7 points
  • Claude Sonnet 4.6: Overall -19 points, code execution -48.4 points, material constraint +16.9 points
  • GPT-o3: Overall -13.9 points, code execution -13.2 points, material constraint -14.7 points

Signals to Watch

  • GPT-o3: Overall plunged 13.9 points
  • GLM-4.6: Material constraint plunged 27.3 points
  • Grok 4: Overall plunged 8.2 points
  • GPT-5.5: Overall plunged 20.6 points
  • Gemini 2.5 Pro: Overall plunged 8.5 points
  • Claude Sonnet 4.6: Overall plunged 19 points

When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness on consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day swings in execution or constraint scores may come from question sampling or may be early signals of genuine degradation, and require follow-up runs to verify.


Data source: YZ Index | Run #256 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!