Claude Sonnet 4.6 and Grok 4 Tie at 96.98: 2026-07-25 Smoke Test Data Brief

On July 25, 2026, the YZ Index Smoke test covered 11 models, with Claude Sonnet 4.6 and Grok 4 both scoring 96.98, tying for first place. Smoke is a daily 10-question quick test designed to observe short-term signals and is not equivalent to the full weekly ranking conclusions.

This Smoke test only covers the two main ranking dimensions of code execution and material constraints. The main ranking formula is 0.55 × Code Execution + 0.45 × Material Constraints. Due to the small daily sample size, single-day scores are more suitable as monitoring signals rather than long-term conclusions about model capabilities.

Daily Rankings

RankModelMain ScoreCode ExecutionMaterial ConstraintsIntegrity
#1Claude Sonnet 4.696.9894.5100pass
#2Grok 496.9894.5100pass
#3DeepSeek V4 Pro92.1694.589.3pass
#4GPT-5.587.1794.578.2pass
#5Gemini 3.1 Pro80.9194.564.3pass
#6GPT-o380.9194.564.3pass
#7Gemini 2.5 Pro80.7394.563.9pass
#8Claude Opus 4.778.4169.589.3pass
#9Doubao Pro71.9869.575pass
#10GLM-4.664.6644.589.3pass
#11Qwen3 Max53.4144.564.3pass

Data Interpretation

Today's top two models, Claude Sonnet 4.6 and Grok 4, both scored 96.98, with code execution at 94.5 and material constraints at 100 for both, showing a balanced high-level combination in both capabilities. DeepSeek V4 Pro ranked third with 92.16, with the same code execution score of 94.5 but a slightly lower material constraints score of 89.3. GPT-5.5 scored 87.17 on the main ranking, with code execution at 94.5 and material constraints at 78.2, reflecting a structural advantage in code execution over material constraints.

Several models showed significant positive changes. Claude Sonnet 4.6 rose 38.3 points on the main ranking, 44.5 points in code execution, and 30.7 points in material constraints. DeepSeek V4 Pro rose 23.5 points on the main ranking, 19.5 points in code execution, and 28.4 points in material constraints. Gemini 3.1 Pro rose 21.4 points on the main ranking and 36.2 points in code execution. These increases come from a single-day small-sample Smoke test and may result from question sampling fluctuations or model performance variations under specific constraints, requiring subsequent runs under the same methodology to confirm stability.

Overall, top models tend to show complementary strengths between code execution and material constraints, while some mid-range models like GLM-4.6 (code execution 44.5, material constraints 89.3) and Qwen3 Max (code execution 44.5, material constraints 64.3) exhibit clear structural differences. All interpretations are based on today's data and do not constitute a judgment on long-term performance.

Key Changes

  • Claude Sonnet 4.6: Main score +38.3, Code Execution +44.5, Material Constraints +30.7
  • DeepSeek V4 Pro: Main score +23.5, Code Execution +19.5, Material Constraints +28.4
  • Gemini 3.1 Pro: Main score +21.4, Code Execution +36.2
  • GPT-o3: Main score +16, Code Execution +19.5, Material Constraints +11.7
  • GPT-5.5: Main score +13.3, Code Execution +19.5, Material Constraints +5.6

Signals to Monitor

  • No abnormal signals were retained for release this time.

When reading Smoke briefs like this, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether the integrity rating has shifted from pass to warn or fail. Large daily fluctuations in execution or constraint scores may result from question sampling or represent early signs of real degradation, requiring subsequent runs for verification.


Data source: YZ Index | Run #245 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!