Grok 4 Tops with 97.66 Points: 2026-09-27 Smoke Quick Test Data Brief

On 2026-09-27, the YZ Index Smoke quick test covered 10 models, with Grok 4 ranking first for the day at 97.66 points. Smoke is a daily 10-question quick test, suitable for observing short-term signals; it is not equivalent to Full weekly ranking conclusions.

This Smoke evaluation covered only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than long-term conclusions about model capabilities.

Daily Ranking

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1Grok 497.6610094.8pass
#2Claude Opus 4.796.019794.8pass
#3GPT-o396.019794.8pass
#4Claude Sonnet 4.692.1993.990.1pass
#5Qwen3 Max86.4479.694.8pass
#6Doubao Pro839272pass
#7GPT-5.578.87287.1pass
#8Gemini 2.5 Pro69.917069.8pass
#9DeepSeek V4 Pro69.3168.969.8pass
#10Gemini 3.1 Pro68.266769.8pass

Data Interpretation

In today's YZ Index Smoke quick test, the top models showed different patterns in the combination of Code Execution and Material Constraints. Grok 4 achieved a main leaderboard score of 97.66 with Code Execution 100 and Material Constraints 94.8, while Claude Opus 4.7 and GPT-o3 both tied at a main leaderboard score of 96.01 with Code Execution 97 and Material Constraints 94.8. Claude Sonnet 4.6 had Code Execution 93.9 and Material Constraints 90.1. Qwen3 Max had Code Execution 79.6 but Material Constraints 94.8, for a main leaderboard score of 86.44; Doubao Pro had Code Execution 92 and Material Constraints 72, for a main leaderboard score of 83, showing a combination of high Code Execution and low Material Constraints.

Compared with the previous run using the same methodology, Grok 4 rose 27.3 points on the main leaderboard, +25 points in Code Execution, and +30 points in Material Constraints; Qwen3 Max rose 11 points on the main leaderboard and +21 points in Material Constraints; GPT-5.5 fell 12.5 points on the main leaderboard and -28 points in Code Execution; Gemini 2.5 Pro fell 10.4 points on the main leaderboard and -17.5 points in Code Execution; Doubao Pro fell 9.8 points on the main leaderboard and -13.7 points in Material Constraints; and DeepSeek V4 Pro fell 8.3 points on the main leaderboard. These movements may stem from question sampling fluctuations or may reflect real changes in single-day performance, and require follow-up runs to verify.

The Smoke quick test is a small-sample single-day signal; GLM-4.6 was not ranked due to incomplete data. The score structure and changes above are drawn directly from that day's data, and no long-term inference is made for now.

Key Changes

  • Grok 4: Main leaderboard up 27.3 points, Code Execution +25 points, Material Constraints +30 points
  • GPT-5.5: Main leaderboard down 12.5 points, Code Execution -28 points, Material Constraints +6.4 points
  • Qwen3 Max: Main leaderboard up 11 points, Material Constraints +21 points
  • Gemini 2.5 Pro: Main leaderboard down 10.4 points, Code Execution -17.5 points
  • Doubao Pro: Main leaderboard down 9.8 points, Code Execution -6.7 points, Material Constraints -13.7 points

Signals to Watch

  • Doubao Pro: Main leaderboard plunges -9.8 points
  • GPT-5.5: Main leaderboard plunges -12.5 points
  • Gemini 2.5 Pro: Main leaderboard plunges -10.4 points
  • DeepSeek V4 Pro: Main leaderboard plunges -8.3 points
  • GLM-4.6: Incomplete data (missing communication dimension, API failure/timeout), has entered automatic rerun, not ranked in this period

When reading this type of Smoke brief, the focus should be on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or they may be early signals of real degradation, and require follow-up runs to verify.


Data source: YZ Index | Run #340 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!