Smoke Evaluation Sees Across-the-Board Plunge: 11 Models Drop 42 Points on Average on Main Leaderboard, Code Execution Dimension Collapses for All

Today's Smoke evaluation results were released at 3 AM, with all 11 mainstream models experiencing a collective collapse on the main leaderboard, averaging a drop of 42 points. Gemini 3.1 Pro topped the list with 40.48 points, but this score itself dropped 33.5 points from yesterday, with only 20 points in the execution dimension and 65.5 points in the constraint dimension.

Why Did the Execution Dimension Suddenly Collapse

The core reason lies in the code execution dimension. Yesterday, most models scored above 100 in execution, but today that figure was directly halved to 20 or 0. Six models—Gemini 3.1 Pro, Doubao Pro, Gemini 2.5 Pro, Grok 4, DeepSeek V4 Pro, and ERNIE Bot 4.5—each scored 20 in execution, while Claude Sonnet 4.6 and the five models below it directly fell to zero.

The formula shows that with an execution weight of 0.55, today's collapse in execution scores directly led to a halving of the overall main leaderboard scores. While the constraint dimension saw minor fluctuations, they were insufficient to offset the execution losses. Qwen3 Max and Claude Opus 4.7 saw their execution scores drop from 100+ to 0, with single-day main leaderboard declines of 52.4 and 52.3 points respectively.

The Real Signals Behind the Rankings

Gemini 3.1 Pro and Doubao Pro tied for the top two spots, both with execution scores of 20 and constraint scores of 65.5 vs. 64.7—a gap of only 0.36 points, indicating that under the current test set, their material constraint capabilities are close, while execution capabilities show no significant differentiation.

Claude Sonnet 4.6 had the highest constraint score of 80.5 in the entire field, yet ranked only 7th due to a zero execution score, confirming a clear disconnect between material constraint and code execution among current models. GPT-5.5 and GPT-o3 both scored 29.93 on the main leaderboard, with constraint scores of 66.5 and execution scores of 0, making it difficult to distinguish between models.

Possible Reasons Behind the Anomaly

An across-the-board crash is extremely rare. The most likely cause is that the difficulty of test questions or evaluation standards were adjusted early this morning. The execution dimension plummeting from high to zero or 20 suggests that newly added questions significantly raised requirements for code correctness, boundary handling, or multi-step reasoning.

Another possibility is that some models experienced server-side degradation or context processing anomalies during the early morning hours, leading to a decline in code execution consistency. Notably, the integrity ratings for Qwen3 Max and Claude Opus 4.7 improved from warn to pass, but their main leaderboard scores still dropped sharply, indicating that integrity improvements cannot compensate for capability gaps.

From an industry perspective, model iterations in May 2026 have entered a refinement stage. After general capabilities converged, code execution has become the most vulnerable dimension. Today's data once again proves that the constraint dimension remains relatively stable, while the execution dimension fluctuates wildly, raising questions about model reliability in real-world engineering scenarios.

When all models fail simultaneously in the same dimension, the problem is likely not with the models themselves, but with the evaluation or the infrastructure.

Today's results send a clear signal to developers: if a task heavily relies on code execution, any current model requires thorough fallback planning and manual verification.


Data source: 赢政指数 (YZ Index) | Run #136 | View raw data