In today's Smoke evaluation, Doubao Pro's material constraint score fell from 80.20 to 64.30, a drop of 15.9 points, while the overall main leaderboard score declined from 91.09 to 83.94.
Score Comparison Data
The code execution dimension held steady at 100.00. The material constraint dimension fell from 80.20 to 64.30. The engineering judgment dimension held steady at 100.00. The task expression dimension fell from 100.00 to 91.70. The integrity rating remained pass.
Breaking Down the Data
The main leaderboard score is formed by weighting two auditable dimensions: code execution and material constraint. Today the material constraint dimension alone dropped 15.9 points, directly dragging the main leaderboard down by 7.2 points. Code execution remained at full marks, indicating the model can still maintain full-mark performance on code-related tasks.
The task expression dimension declined by 8.3 points, falling at the same time as the material constraint dimension, indicating that the model showed volatility in instruction following and output format control on the questions drawn for this round.
Possible Causes
The Smoke evaluation includes only 10 questions per day, 2 per dimension, so the randomness of question drawing may cause single-day score fluctuations. This 15.9-point drop in the material constraint dimension exceeds the typical range of draw-related fluctuation, so a distinction must be made between changes in question difficulty and genuine degradation of the model's capabilities.
Based on the available scores, code execution and engineering judgment both remain at full marks, indicating no systemic decline in the model's underlying capabilities. The simultaneous impact on material constraint and task expression is more likely related to the characteristics of the material-fidelity questions drawn for this round.
Implications for Users
Teams that rely heavily on code execution can continue to use Doubao Pro; the code execution dimension remains at 100.00 and was unaffected by this fluctuation.
For scenarios sensitive to material fidelity — such as document summarization, contract extraction, and knowledge base question answering — an additional manual verification step is needed. The drop in material constraint from 80.20 to 64.30 means the model was more prone in this test to producing information not present in the source material or omitting key constraints.
The task expression dimension fell to 91.70, meaning automated pipelines that depend on strict format output may encounter formatting errors; adding an output validation step in production environments is recommended.
Strategic Assessment
A single-day drop of 15.9 points in the material constraint dimension exceeds the normal draw-fluctuation range and warrants continued tracking of the same dimension in the next Smoke evaluation. If the material constraint score rebounds above 75 in the next round, this decline is more likely to have been caused by question drawing; if it remains below 70, the model's actual stability in the area of material constraint must be reconsidered.
Code execution remains at full marks, indicating Doubao Pro is still competitive in that direction. At 83.94, the main leaderboard score is already below yesterday's level, so when selecting a model, material constraint should be treated as an independent evaluation item rather than relying solely on the main leaderboard ranking.
The integrity rating remains pass; the model showed no integrity issues in this test and can remain in the candidate pool, but material constraint performance should be a key validation metric.
Data source: YZ Index | Run #324 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接