Doubao Pro scored 83.00 on today's Smoke benchmark main leaderboard, down 9.8 points from yesterday's 92.85, with the material constraint dimension falling from 85.70 to 72.00.
Breaking Down the Data
Code execution fell from 98.70 to 92.00, a drop of 6.7 points; material constraint fell by 13.7 points; engineering judgment dropped from 100.00 to 75.00, a decline of 25 points; task expression rose from 91.70 to 100.00, up 8.3 points. The main leaderboard is weighted only by code execution and material constraint, so the sharp slide in material constraint directly dragged down the overall ranking.
Cause Analysis: Sampling Noise or Real Degradation
The Smoke benchmark contains only 10 questions per day, 2 per dimension—an extremely small sample, where a change in a single question's score can produce swings of more than 10 points. The simultaneous marked declines in material constraint and engineering judgment most likely stem from today drawing questions that demand stricter instruction following and multi-step reasoning. The rise in task expression likewise points to randomness in question distribution rather than an overall degradation of model capability. The integrity rating remains pass, with no consistency issues observed.
The combination of 72.00 in material constraint and 75.00 in engineering judgment points to sampling-related deviation in the model's multi-step decision stability under strict material constraints.
What This Means for Users
Developers of RAG scenarios that rely heavily on material constraints should note that today's 72.00 means a higher risk of hallucination or omission in tasks with strict requirements to cite the source text. Code execution still holds at 92.00, so the impact on pure code-generation tasks is limited. The 75.00 in engineering judgment suggests that for complex prompt scenarios requiring engineering trade-offs, adding a manual review step is advisable.
- Teams sensitive to material constraints: today's score has fallen below the 75-point threshold, so output validation in production environments should be strengthened.
- Teams focused mainly on code execution: 92.00 remains within the usable range, so continued use is fine.
Strategic Assessment
Based on a single day's data, Doubao Pro's 9.8-point drop on the main leaderboard is driven mainly by question sampling variance in the material constraint and engineering judgment dimensions, not by genuine model degradation. If material constraint remains below 80 in the next Smoke benchmark run, whether a systemic problem has emerged will need to be verified with priority. Current data supports only the conclusion that "sampling variance is the main factor"; continued observation of the standard deviation for the same dimension across three consecutive days is recommended.
Data source: Winzheng (YZ Index) | Run #340 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接