Doubao Pro Main Leaderboard Plummets 12.6 Points, Code Execution Drops 25 Points in a Single Day

Doubao Pro scored 82.16 on the main leaderboard in today's Smoke evaluation, down from 94.74, with the code execution dimension falling from 100.00 to 75.00.

Score Comparison Data

The specific changes from yesterday to today are: code execution 100.00→75.00 (-25 points), material constraints 88.30→90.90 (+2.6 points), engineering judgment 75.00→75.00 (0 points), task expression 91.70→90.00 (-1.7 points), and overall main leaderboard 94.74→82.16 (-12.6 points). Integrity rating remains pass.

Data Fact Breakdown

The Smoke evaluation covers only 2 questions per dimension per day, 10 questions in total. The single-day loss of 25 points on the code execution dimension directly dragged the main leaderboard down by 12.6 points. Material constraints actually rose by 2.6 points, indicating improved performance on the two questions in that dimension. The changes in engineering judgment and task expression were both within 2 points, which falls within the normal range.

Cause Analysis

This decline is most likely attributable to question sampling fluctuation. The Smoke evaluation has an extremely small sample size; differences in difficulty or specific requirements between the two code execution questions alone could create a 25-point gap. The increase in material constraints suggests the model has not undergone systematic degradation. The zero change in engineering judgment further supports the conclusion of fluctuation rather than capability decline.

Implications for Users

Teams that rely heavily on code execution should note that single-day Smoke results may amplify sampling error, and it is recommended to observe results over multiple consecutive days before making model selection decisions. For scenarios sensitive to material fidelity, Doubao Pro's score of 90.90 today is actually better than yesterday's, which can serve as a reference basis.

Strategic Assessment

The single-day 12.6-point drop on the main leaderboard is primarily determined by the two code execution questions and does not constitute sufficient evidence of genuine model regression. The signal to verify in the next evaluation is: if code execution scores below 85 for two consecutive days, its code capability stability needs to be reassessed. The current data only shows normal sampling fluctuation.

Taking all dimensions into account, Doubao Pro's performance today remains within the usable range, with an integrity rating of pass and no access warnings triggered. Developers can continue to use it, but should expand the test sample for code execution tasks to at least 10 or more questions to reduce the impact of single-day fluctuation.


Data source: YZ Index | Run #283 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!