In the YZ Index Smoke Lightweight Evaluation on June 24, 2026, ERNIE Bot 4.5's main leaderboard score plunged 34.1 points to 64.63 compared to the previous day, with the execution dimension dropping directly from 100 to 50.
Obvious Disparity Between Execution and Constraint
Today, the top three on the main leaderboard—DeepSeek V4 Pro, Gemini 3.1 Pro, and Grok 4—all achieved code execution scores of 100 and material constraint scores of 100. The fourth to sixth places—Doubao Pro, Gemini 2.5 Pro, and GPT-5.5—maintained execution scores of 100 and constraint scores of 94.5, with main leaderboard scores all at 97.53.
The eighth place, Claude Opus 4.7, and the ninth place, Qwen3 Max, both have a main leaderboard score of 72.5, with execution scores of 50 and constraint scores of 100. The tenth place, Claude Sonnet 4.6, has an execution score of 50, a constraint score of 95.5, and a main leaderboard score of 70.48. This combination of an execution score of 50 and a constraint score near perfect constitutes the typical structure of the lower half of today's leaderboard.
Execution Scores of Four Models Halved Collectively
Comparison with yesterday shows that ERNIE Bot 4.5's execution score dropped by 50 points, Claude Opus 4.7's execution score dropped by 50 points, Claude Sonnet 4.6's execution score dropped by 50 points, and Qwen3 Max's execution score dropped by 50 points. The execution dimension of all four models experienced a 50-point cliff drop, causing their main leaderboard scores to fall by 34.1, 27.5, 24.4, and 1.5 points respectively.
Changes in the material constraint dimension were relatively moderate. Claude Sonnet 4.6's constraint score actually rose by 6.9 points to 95.5, while ERNIE Bot 4.5's constraint score dropped by 14.7 points to 82.5 and received a warn rating. The constraint changes for the remaining models did not exceed 10 points.
Score Structure Reveals Capability Boundaries
The top seven models all maintained an execution dimension score of 100, with constraint dimension scores fluctuating between 94 and 100, indicating that these models maintain stable output on code execution tasks. The eighth to eleventh models collectively stalled at 50 points in the execution dimension, yet their constraint dimensions reached 82.5-100 points, indicating that constraint tasks put significantly less pressure on these models than execution tasks.
In the core_overall formula, the weight for code execution is 0.55, higher than the 0.45 for material constraint. Therefore, a drop in execution dimension from 100 to 50 has a greater direct impact on the total main leaderboard score than an equivalent change in the constraint dimension, which is fully consistent with the magnitude of today's declines in the four models.
The combination of an execution score of 50 and a constraint score of 100 has become a fixed pattern in the lower half of today's leaderboard.
ERNIE Bot 4.5 simultaneously exhibited a warn signal and the largest drop, indicating significant fluctuations in both execution and constraint dimensions. The other three models with sharp execution drops still maintain a pass rating, indicating that the integrity dimension has not triggered a new threshold.
Today's data only reflects the results of a single 10-question rapid test. The large fluctuations in the execution dimension may stem from differences in question difficulty distribution or model output stability in this instance. Subsequent multi-day data is needed to verify whether a sustained trend is forming.
Data source: YZ Index | Run #195 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接