In today's Smoke evaluation, Doubao Pro's material constraint score dropped from 100.00 to 79.50, a decline of 20.5 points, while its main leaderboard score rose from 86.25 to 90.78.
Score Change Details
Code execution rose from 75.00 to 100.00, an increase of 25 points. Engineering judgment fell from 100.00 to 75.00, a decline of 25 points. Task expression rose from 91.70 to 95.00, an increase of 3.3 points. Integrity rating remained "pass."
Data Fact Breakdown
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension. Material constraint lost 20.5 points in a single day, and engineering judgment lost 25 points, together dragging down side leaderboard performance. However, the perfect score in code execution directly boosted the main leaderboard by 4.5 points, showing a clear offset effect across dimensions.
Potential Cause Analysis
Question-drawing variance is the primary explanation. If the two daily material constraint questions involve higher difficulty or stricter citation requirements, the model's fidelity losses would be amplified. Engineering judgment also dropped 25 points, suggesting that the day's questions may have simultaneously tested scenarios combining multi-step reasoning with constraint adherence.
The possibility of genuine model degradation is low. Code execution jumped directly from 75.00 to a perfect score, indicating no systemic decline in core capabilities. Task expression actually rose slightly, and the overall main leaderboard score remained positive.
Implications for Users
Teams prioritizing code execution can continue using Doubao Pro; today's 100.00 score shows it has reached a perfect level in that dimension. For scenarios sensitive to material fidelity, such as contract extraction, policy citation, and academic writing, the 79.50 score level means additional manual verification steps are needed.
The 75.00 engineering judgment score suggests that in complex decision-chain tasks, the model may exhibit judgment deviations under constraints. Developers relying on this capability should prepare fallback options or secondary verification processes.
Strategic Assessment
The simultaneous single-day decline in material constraint and engineering judgment, combined with the sharp rise in code execution, points more toward question-drawing variance than model capability degradation. The main leaderboard rise masks side leaderboard issues; if material constraint remains below 85 points in the next round, consistency verification will be a priority.
Current data does not support a "persistent degradation" conclusion, but the simultaneous 25-point drop in engineering judgment and 20.5-point drop in material constraint are signals worth tracking in the next round. It is recommended to monitor the standard deviation of material constraint scores across three consecutive Smoke evaluations; if it exceeds 15 points, consider adjusting usage strategy.
Overall, Doubao Pro's performance today shows clear dimensional divergence. The main leaderboard benefited from a perfect code execution score, while the declines in material constraint and engineering judgment are more likely attributable to question drawing. Users in material-constraint-heavy scenarios should remain vigilant and continuously track subsequent scores to confirm whether this is random fluctuation.
Data source: YZ Index (YZ Index) | Run #297 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接