Doubao Pro Smoke Evaluation Shows All 5 Dimensions Missing in 0-Score Anomaly, with API Timeout as Main Cause

Doubao Pro exhibited completely missing data across all five dimensions—execution, grounding, judgment, integrity, and communication—in today's Smoke evaluation, and its main leaderboard ranking was directly canceled.

Data Facts: All Five Dimensions Simultaneously Scoreless

The score comparison shows that code execution, material constraints, engineering judgment, task expression, and the main leaderboard are all marked as “- → - (-)”. The Smoke evaluation runs a fixed set of 10 questions per day (2 per dimension). In this run, API responses timed out, resulting in no valid scores being generated for any auditable dimension or side leaderboard dimension.

Cause Analysis: API Timeout Rather Than Question Fluctuation

A single-day 10-question quick test inherently involves sampling fluctuation, but the simultaneous absence of all five dimensions exceeds the normal range. Normal fluctuation typically appears as score variation on individual questions, not an interface-level timeout. The provided comparison data only records API failures/timeouts and shows no question response records at all, so the most likely cause is an interruption in the call chain, rather than degradation of the model's ability to answer specific questions.

The simultaneous absence of scores in execution and material constraints—the two main leaderboard dimensions—points to underlying API response failure, rather than a localized issue with engineering judgment or task expression.

Implications for Users: Teams Relying on Code Execution Should Switch Immediately

For teams that depend on Doubao Pro for code execution, this evaluation yielded no execution score, meaning they cannot verify the model's day-to-day consistency through Smoke results. Scenarios sensitive to material constraints likewise lose their reference basis. Developers should temporarily migrate critical tasks to other models that have completed evaluations before rerun results are available, to avoid production interruptions caused by API instability.

  • Code execution scenarios: Today's execution accuracy cannot be confirmed; it is recommended to pause automated script calls.
  • Material constraint scenarios: Long-text or instruction-following tasks should be switched to other available models.

Strategic Assessment: This Anomaly Does Not Constitute a Degradation Signal

Based on the existing score comparison, Doubao Pro's missing scores this round represent an interface-level event rather than a decline in model capability. The stability dimension measures the standard deviation of scores across multiple responses to similar questions; this time, with no scores available, it cannot be calculated, so the absence should not be interpreted as reduced stability. If the next Smoke rerun restores normal scores, this event can be treated as a one-off occurrence; if the rerun still shows missing scores, API availability warrants close attention.

For teams currently in the model selection process<|eos|>


Data Source: YZ Index | Run #260 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!