Grok 4 Leads with 98.41 Points: 2026-08-12 Smoke Quick Test Data Briefing

The 2026-08-12 YZ Index Smoke quick test covered 11 models, with Grok 4 ranking first that day with 98.41 points. Smoke is a daily 10-question quick test suited for observing short-term signals, and is not equivalent to the Full weekly ranking conclusions.

This Smoke evaluation only covers two main leaderboard dimensions — Code Execution and Material Constraint — with the main leaderboard formula being 0.55 × Code Execution + 0.45 × Material Constraint. Due to the small daily sample size, single-day scores are better suited as monitoring signals rather than long-term conclusions about model capability.

Daily Rankings

RankModelMain ScoreCode ExecutionMaterial ConstraintIntegrity
#1Grok 498.4197.1100pass
#2Gemini 2.5 Pro88.5399.675pass
#3Claude Opus 4.786.0374.6100pass
#4GLM-4.679.3862.5100pass
#5Claude Sonnet 4.6757575pass
#6Gemini 3.1 Pro757575pass
#7GPT-o3757575pass
#8Doubao Pro74.6274.375pass
#9DeepSeek V4 Pro70.1274.365pass
#10Qwen3 Max62.7173.150pass
#11GPT-5.561.255075pass

Data Analysis

In today's YZ Index Smoke quick test, the leading models showed clear differentiation in how they combined Code Execution and Material Constraint scores. Grok 4 ranked first with a main score of 98.41, with Code Execution at 97.1 and Material Constraint at 100, maintaining high performance on both capabilities. Gemini 2.5 Pro scored 88.53 on the main leaderboard, with Code Execution at 99.6 but Material Constraint at only 75, highlighting its strength in the Code Execution dimension. Claude Opus 4.7 scored 86.03 on the main leaderboard, with Material Constraint at 100 and Code Execution at 74.6, reflecting a relative advantage in Material Constraint. GLM-4.6 scored 79.38 on the main leaderboard, with Code Execution at 62.5 and Material Constraint at 100, similarly relying on its Material Constraint score to support overall performance.

Multiple models showed notable score changes. GPT-5.5 dropped 21.9 points on the main leaderboard, with Code Execution down 50 points and Material Constraint up 12.4 points; Gemini 3.1 Pro dropped 19.7 points on the main leaderboard, with Code Execution down 25 points and Material Constraint down 13.3 points; Qwen3 Max dropped 18.9 points on the main leaderboard, with Code Execution down 14.4 points and Material Constraint down 24.3 points; DeepSeek V4 Pro dropped 17.9 points on the main leaderboard, with Code Execution down 25.7 points and Material Constraint down 8.3 points; Claude Sonnet 4.6 dropped 10.7 points on the main leaderboard, with Code Execution down 25 points and Material Constraint up 6.7 points. These single-day fluctuations may stem from differences in question sampling, or may reflect instability in specific model capabilities, and require confirmation through subsequent runs using the same methodology.

Regarding anomaly signals, GLM-4.6's integrity rating dropped to Fail, with the previous record showing fail→pass. This change appeared in a small-sample single-day test, and the specific cause still requires more run data to verify, in order to distinguish random fluctuations from genuine regression. The overall analysis is based solely on the day's score structure, avoiding conclusions about long-term model performance.

Key Changes

  • GPT-5.5: Main leaderboard down 21.9 points, Code Execution -50 points, Material Constraint +12.4 points
  • Gemini 3.1 Pro: Main leaderboard down 19.7 points, Code Execution -25 points, Material Constraint -13.3 points
  • Qwen3 Max: Main leaderboard down 18.9 points, Code Execution -14.4 points, Material Constraint -24.3 points
  • DeepSeek V4 Pro: Main leaderboard down 17.9 points, Code Execution -25.7 points, Material Constraint -8.3 points
  • Claude Sonnet 4.6: Main leaderboard down 10.7 points, Code Execution -25 points, Material Constraint +6.7 points

Signals to Watch

  • GLM-4.6: Integrity rating dropped to Fail (fail→pass)

When reading this type of Smoke briefing, the focus should be on two questions: first, whether a model has exposed the same type of weakness on multiple consecutive days; second, whether the integrity rating has moved from pass to warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be early signals of genuine regression, requiring follow-up runs for verification.


Data source: YZ Index | Run #275 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!