Doubao Pro Material Constraint Drops 40.9 Points in a Single Day; Code Execution Up 25 Points; Main Leaderboard Slips 4.7

Doubao Pro's material constraint score in today's Smoke evaluation dropped from 90.90 to 50.00, while its code execution score rose from 75.00 to 100.00, causing the main leaderboard to slip overall from 82.16 to 77.50.

Data Fact Breakdown

Today's Smoke evaluation contains only 10 questions, with 2 per dimension. The material constraint dimension scored 50.00 today, down 40.9 points from yesterday's 90.90; the code execution dimension scored 100.00 today, up 25 points from yesterday's 75.00; engineering judgment held steady at 75.00; task expression edged up to 91.70. Since the main leaderboard is weighted solely by code execution and material constraint, the sharp drop in material constraint directly pulled down the overall ranking.

Cause Analysis: Lottery Fluctuation or Capability Degradation

Smoke evaluation draws questions randomly each day, and with a sample size of just 2 questions, a single misstep can cause dramatic swings in dimension scores. Yesterday's material constraint score of 90.90 reflected high material fidelity, while today's 50.00 suggests the model may have significantly deviated or omitted content in the two source-alignment questions. The simultaneous rise of code execution to a perfect score indicates that today's code questions happened to match the model's current capabilities in difficulty or type. Engineering judgment and task expression changed only marginally, further confirming that the fluctuation is concentrated in the auditable material constraint dimension.

If this were genuine degradation, it would typically be accompanied by synchronized declines across multiple dimensions or changes in integrity ratings. However, today's integrity rating remains "pass," and engineering judgment has not moved either. Therefore, the 40.9-point drop in material constraint is more likely attributable to sample variance from question selection than to parameter-level model degradation.

Specific Implications for Users

For RAG retrieval, enterprise document Q&A, and contract review scenarios that rely heavily on material constraint, today's data suggests adding manual review steps. The perfect code execution score indicates stable performance on pure programming tasks today, making the model suitable for ad-hoc code generation needs. For hybrid workflows requiring both source alignment and code capabilities, splitting model calls or adding multi-round validation is recommended.

Strategic Judgment

Based on the single-day comparison, Doubao Pro's 4.7-point main leaderboard drop is primarily driven by material constraint fluctuation, with no evidence currently supporting sustained degradation. If material constraint remains below 70 in the next Smoke evaluation, the priority level should be raised; if it rebounds above 80, the drop can be deemed an occasional random-selection event. The current data does not support downgrading the model's overall capability assessment, but teams that depend on material constraint are reminded to keep multi-model backups for critical projects.


Data source: YZ Index | Run #284 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!