Grok 4 Leads with 87 Points: 2026-09-13 Smoke Quick-Test Data Brief

The 2026-09-13 YZ Index Smoke quick test covered 10 models, with Grok 4 ranking first for the day at 87 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to the conclusions of the weekly Full leaderboard.

This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The Main Leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Grok 48710071.1pass
#2Doubao Pro85.1510067pass
#3GPT-o385.1510067pass
#4DeepSeek V4 Pro77.347580.2pass
#5Gemini 3.1 Pro73.257571.1pass
#6Claude Opus 4.763.595080.2pass
#7GPT-5.560.877543.6pass
#8Gemini 2.5 Pro59.55071.1pass
#9Qwen3 Max57.655067pass
#10Claude Sonnet 4.653.565057.9pass

Data Interpretation

Today's YZ Index Smoke quick test shows clear divergence among leading models in how Code Execution and Material Constraints pair together. Grok 4 has a Main Leaderboard score of 87, with Code Execution 100 and Material Constraints 71.1; Doubao Pro and GPT-o3 both have Main Leaderboard scores of 85.15, both with Code Execution 100 and Material Constraints 67. All three occupy top positions on the strength of high Code Execution scores. DeepSeek V4 Pro has a Main Leaderboard score of 77.34, with Code Execution 75 and Material Constraints 80.2; Gemini 3.1 Pro has a Main Leaderboard score of 73.25, with Code Execution 75 and Material Constraints 71.1. The two have relatively stronger Material Constraints, forming a different balance structure from the top three.

Several models showed significant changes: Qwen3 Max fell 25.9 points on the Main Leaderboard and 49.2 points in Code Execution; Claude Sonnet 4.6 fell 22.6 points on the Main Leaderboard and 45.5 points in Code Execution, while Material Constraints rose 5.3 points; GPT-5.5 fell 21.4 points on the Main Leaderboard, 22 points in Code Execution, and 20.7 points in Material Constraints; Claude Opus 4.7 fell 18.7 points on the Main Leaderboard and 47 points in Code Execution, while Material Constraints rose 15.9 points; Gemini 2.5 Pro fell 16.6 points on the Main Leaderboard, 23.5 points in Code Execution, and 8.2 points in Material Constraints. These anomalous signals may stem from question sampling fluctuations, or they may reflect genuine single-day performance regression, and require follow-up run review for confirmation.

GLM-4.6's data was incomplete due to API failures/timeouts, with several missing evaluation dimensions, so it did not participate in this edition's ranking. As a small-sample, single-day signal, the observations above only reflect the day's score structure, and the wording remains measured.

Key Changes

  • Qwen3 Max: Main Leaderboard down 25.9 points, Code Execution -49.2 points
  • Claude Sonnet 4.6: Main Leaderboard down 22.6 points, Code Execution -45.5 points, Material Constraints +5.3 points
  • GPT-5.5: Main Leaderboard down 21.4 points, Code Execution -22 points, Material Constraints -20.7 points
  • Claude Opus 4.7: Main Leaderboard down 18.7 points, Code Execution -47 points, Material Constraints +15.9 points
  • Gemini 2.5 Pro: Main Leaderboard down 16.6 points, Code Execution -23.5 points, Material Constraints -8.2 points

Signals to Watch

  • Claude Opus 4.7: Main Leaderboard plunged -18.7 points
  • GPT-5.5: Main Leaderboard plunged -21.4 points
  • Gemini 2.5 Pro: Main Leaderboard plunged -16.6 points
  • Qwen3 Max: Main Leaderboard plunged -25.9 points
  • Claude Sonnet 4.6: Main Leaderboard plunged -22.6 points
  • GLM-4.6: Data incomplete (missing multiple evaluation dimensions due to API failure/timeout); has been queued for automatic rerun and is not included in this edition's ranking

When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its Integrity rating has moved from pass to warn or fail. Large single-day changes in Code Execution or Material Constraints scores may come from question sampling, or they may be early signals of genuine regression, requiring follow-up run review.


Data: YZ Index | Run #320 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!