In the 2026-06-23 Smoke lightweight evaluation, Qwen3 Max scored 74 on the main leaderboard, with 100 in Execution, 95.7 in Material Constraint, and Integrity directly failed, dropping 12 points from yesterday's main leaderboard, becoming the only model below 80 among the 11 models.
Perfect-Score Models Coexist with Constraint Shortcomings
Three models—Claude Opus 4.7, Gemini 3.1 Pro, and Grok 4—each scored 100 on the main leaderboard, with perfect 100 in both Execution and Material Constraint, all Integrity pass, forming the only combination without any shortcomings. DeepSeek V4 Pro followed closely with a main leaderboard score of 99.37, Execution 100, Constraint 98.6, also pass.
ERNIE Bot 4.5 scored 98.74 on the main leaderboard, with Execution 100, Constraint 97.2, and Integrity warn. Doubao Pro scored 98.07 on the main leaderboard, with Execution 100, Constraint 95.7, pass. GPT-o3 scored 96.81 on the main leaderboard, with Execution 100, Constraint 92.9, pass. Gemini 2.5 Pro and GPT-5.5 tied at 96.18 on the main leaderboard, both with Execution 100 and Constraint 91.5, pass. Claude Sonnet 4.6 scored 94.87 on the main leaderboard, with Execution 100, Constraint 88.6, pass.
Execution Dimension Uniform, Constraint Determines Ranking
All 11 models scored a perfect 100 in the Execution dimension, making the 0.55-weight portion of the formula identical, leaving Material Constraint with 0.45 weight as the sole ranking criterion. Constraint scores range from 100 to 88.6, and Qwen3 Max's 95.7 is heavily penalized by its Integrity fail, showing the direct penalty of Integrity rating on the final main leaderboard.
In yesterday's comparison, ERNIE Bot 4.5's main leaderboard rose by 50.8 points, with Constraint recovering 51.7 points from a low; Gemini 2.5 Pro's main leaderboard rose by 24.9 points, with Constraint changing by -5.9 points; Qwen3 Max's Constraint increased by 26.9 points, but its main leaderboard still dropped 12 points, indicating that the Constraint improvement was insufficient to offset the deduction caused by Integrity fail.
Anomalous Signals Point to Integrity Threshold
Qwen3 Max is the only model with Integrity fail, scoring 74 on the main leaderboard—far below the theoretical value corresponding to its Constraint score of 95.7—demonstrating that the fail directly triggered an additional penalty mechanism. Claude Sonnet 4.6, with Constraint 88.6, still maintains pass and a main leaderboard score of 94.87, showing the protective effect of the pass threshold on rankings.
Perfect Execution, differentiated Constraints, and Integrity veto power are the clearest structural characteristics of the scores in this Smoke evaluation.
When Execution is no longer a differentiator, Material Constraint and Integrity rating together determine who stays in the first tier.
Data source: YZ Index | Run #194 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接