In today's Smoke evaluation, Doubao Pro's code execution dimension score dropped from 91.70 yesterday to 75.00, a decrease of 16.7 points; the material constraint dimension rose from 69.40 to 100.00, an increase of 30.6 points; the main leaderboard score rose from 81.67 to 86.25, an increase of 4.6 points.
Direct Facts from Score Comparison
The engineering judgment dimension rose from 63.90 to 100.00, an increase of 36.1 points; the task expression dimension fell from 100.00 to 91.70, a decrease of 8.3 points. The integrity rating remains pass. The Smoke evaluation includes only 10 questions per day, 2 questions per dimension, so single-day fluctuations are within the normal range.
Cause Analysis: Question Sampling Fluctuation Rather Than Genuine Model Degradation
The two main leaderboard dimensions, code execution and material constraint, showed dramatic changes in completely opposite directions. While code execution dropped 16.7 points, material constraint rose 30.6 points, and engineering judgment also increased by 36.1 points in tandem. This opposing trend most likely stems from differences in difficulty and question types in the daily question sampling. The Smoke evaluation question bank is limited, and the random combination of 10 questions per day can easily lead to the model encountering harder code execution cases in one dimension while facing easier scoring cases in material constraint and engineering judgment.
If this were genuine model capability degradation, it would typically be accompanied by simultaneous or same-direction declines across multiple dimensions, rather than substantial improvements in both engineering judgment and material constraint at the same time. Task expression only declined slightly by 8.3 points, further confirming that overall capability has not experienced a systematic decline. The main leaderboard score actually rose by 4.6 points, which also contrasts with the single-dimension plunge in code execution, indicating that the full-score boost from material constraint outweighed the losses from code execution.
Specific Implications for Users
Development teams that heavily rely on code execution should note that today's 75.00 score indicates a noticeable decline in Doubao Pro's completion rate or accuracy on code tasks similar to those in the Smoke evaluation. It is advisable to add manual review steps for critical code generation tasks, especially in scenarios requiring multi-step debugging or complex algorithm implementation.
For scenarios with high requirements on material fidelity, today's 100.00 score provides a positive signal, as the model achieved a perfect score in following given materials and avoiding hallucinated output. The 100.00 engineering judgment score likewise shows that today's sampled questions allowed Doubao Pro to perform excellently on decision-making problems requiring engineering trade-offs.
Strategic Assessment
Based on today's score comparison, Doubao Pro's Smoke evaluation results are primarily driven by question sampling fluctuation rather than genuine model degradation. The fact that the main leaderboard score rose by 4.6 points shows that the single-dimension plunge did not change its overall ranking position. The signal that requires verification in the next round is: if the code execution dimension remains around 75 points for several consecutive days, the model's stability in that dimension needs to be reassessed.
Current data does not support interpreting this plunge as a permanent capability decline. Enterprises in the selection process can continue monitoring the paired trends of code execution and material constraint in future Smoke evaluations to confirm whether the fluctuation converges.
Data source: YZ Index | Run #295 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接