Three Models Tie for First Place in Smoke Ranking, Full Score on Execution but Constraint Warnings

Smoke's quick test results today show that Claude Opus 4.7, Claude Sonnet 4.6, and GPT-5.5 all scored 87.76 on the main ranking, tying for first place. The core reason is that all three achieved a perfect 100 on the code execution dimension, while scoring 72.8 on material constraint, triggering a warn signal.

Perfect Execution Has Become the Standard; Constraint Is the Only Differentiator

All top eight models scored 100 on code execution, indicating that current mainstream models have reached saturation on simple coding tasks. The real gap is only in material constraint. The 72.8 score of Claude and GPT-5.5 leads Doubao Pro's 70.8 and Gemini 2.5 Pro's 70, a small yet decisive difference that determines the top three rankings.

The material constraint dimension primarily evaluates a model's fidelity to given materials and its boundary control. The warn rating at 72.8 means these models exhibited slight over-inference or information spillover in some questions. In contrast, DeepSeek V4 Pro and Grok 4 triggered a fail on the constraint dimension, directly dropping to 9th and 10th place on the main ranking.

ERNIE Bot Execution Collapse Creates a Clear Gap

ERNIE Bot 4.5 scored only 50 on execution, with an overall main ranking of 56.3, clearly at the bottom. This model can no longer compete with mainstream models in code execution, exposing its long-standing weakness in engineering tasks.

No model showed significant fluctuations today; all models' scores are consistent with yesterday's, so the stability dimension yields no new signals. On the industry front, the Claude series and GPT-5.5 hit the same score on the constraint dimension, suggesting that the current training paradigm for improving "material boundary control" has entered a bottleneck phase.

A perfect 100 on execution is merely the baseline; a warn on constraint is the true ceiling.

In the short term, model iteration efforts will continue to focus on refining material constraint; otherwise, even higher execution scores cannot push the main ranking upward overall.


Data source: YZ Index | Run #145 | View raw data