In today's Smoke evaluation, Doubao Pro's material constraint score fell from 85.90 yesterday to 58.30, a drop of 27.6 points. Meanwhile, its code execution score climbed from 50.00 to 99.30, a gain of 49.3 points, and its main leaderboard score rose from 66.16 to 80.85.
Data Breakdown
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension. Material constraint and code execution constitute the main leaderboard, while engineering judgment and task expression form the side leaderboard. Today's scores: material constraint 58.30, code execution 99.30, and main leaderboard 80.85. Yesterday's scores: material constraint 85.90, code execution 50.00, and main leaderboard 66.16. Engineering judgment fell from 100.00 to 44.50, while task expression rose from 91.70 to 95.00. The integrity rating remains pass.
Cause Analysis: Sampling Fluctuation or Model Degradation
With just 2 questions per day in the material constraint dimension, a single-question misstep can produce swings of 20-30 points. Today's material constraint score of 58.30 likely corresponds to cases in which multiple questions failed to strictly adhere to the given materials, whereas yesterday's 85.90 reflects relatively high material fidelity. The jump in code execution from 50.00 to 99.30 likewise points to differences in daily question sampling: yesterday's two questions may have involved complex debugging or edge-case errors, while today's two may have been standard implementation questions. The 55.5-point drop in engineering judgment (side leaderboard, AI-assisted evaluation) further confirms the amplifying effect that a single day's question set has on side leaderboard dimensions. Taken together, the sharp opposing movements in the two core dimensions are most likely caused by question sampling fluctuation rather than genuine degradation at the model-parameter level.
Implications for Users
Teams that emphasize code execution may achieve higher pass rates under today's data, but scenarios sensitive to material fidelity, such as contract extraction and policy interpretation, require extra manual checks. Prototype validation workflows that rely on engineering judgment (side leaderboard, AI-assisted evaluation) may face higher rework risk. The task expression score of 95.00 is relatively stable, so copywriting scenarios with strict output-format requirements will see limited impact. Enterprises selecting models should treat today's main leaderboard score of 80.85 as a high-volatility sample rather than a stable capability benchmark.
Strategic Assessment
The main leaderboard score of 80.85 was driven by code execution at 99.30, while material constraint at 58.30 has fallen below yesterday's level. A single-day opposing swing of more than 45 points indicates that the current test sample size is insufficient to distinguish sampling noise from genuine capability change. If material constraint continues to score below 70 in the next Smoke evaluation, it should be given higher priority; if code execution retreats while material constraint recovers, today's data is more likely an extreme sampling outcome. Based on the current comparison, Doubao Pro's main leaderboard improvement comes mainly from the code execution dimension, and continuous tracking of material constraint will be the key signal for judging its stability.
Engineering judgment (side leaderboard, AI-assisted evaluation) falling to 44.50 at the same time as material constraint dropped to 58.30 suggests that the model may lack sufficient consistency on composite tasks requiring both strict constraints and engineering decision-making. Developers building workflows that demand multiple rounds of material verification should reserve manual review checkpoints. Although the main leaderboard score of 80.85 is higher than yesterday's, the 41-point gap between the two dimensions that comprise it shows that a single main leaderboard figure masks internal imbalance.
The integrity rating remains pass, indicating that the model has not exhibited systematic hallucination or refusal to answer questions. The stability dimension was not provided in this round of data, so score standard deviation cannot be directly assessed. Users can calculate the standard deviation of material constraint scores from three consecutive days of Smoke data to determine whether the model has entered a high-volatility range.
Overall, the most direct interpretation of today's data is that sampling fluctuation dominates. The simultaneous plunge in material constraint and surge in code execution do not yet constitute sufficient evidence of genuine model degradation, but they do form a signal that requires validation by the next round of data. Teams in material-constraint-heavy scenarios should hold off on expanding deployment and wait for material constraint scores to stabilize.
Data source: YZ Index | Run #312 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接