Doubao Pro's Smoke evaluation today saw complete data loss across five dimensions—execution, material constraints, judgment, integrity, and communication—due to an API failure/timeout. It has entered the automatic re-run process and will not participate in the current main leaderboard ranking.
Data Facts: All Five Dimensions Show "-"
The score comparison table shows that code execution, material constraints, engineering judgment, task expression, and the main leaderboard all show "-" to "-", with no numerical output. This directly conflicts with the Smoke evaluation's normal routine of 10 questions per day (2 per dimension), indicating that the request was interrupted during the execution phase.
Cause Analysis: Clearly Points to API-Level Failure
Fluctuations in question selection typically only cause individual question scores to rise or fall, not entire dimensions to be recorded as empty. A genuine model degradation would manifest as existing scores with lower values, whereas in this record, even the scores themselves do not exist. The system explicitly indicates "API failure/timeout," which points to network connectivity, authentication, or server-side response timeouts rather than any change in the model's capabilities.
As a daily quick test, single-day fluctuations in the Smoke evaluation are normal, but complete dimensional data loss is a different kind of signal. It rules out the possibilities of "bad luck in question selection" or "sudden model weakening," pointing directly to an interruption in the external call chain.
Implications for Users
Teams that heavily rely on code execution should add API health checks and retry mechanisms when integrating Doubao Pro, to prevent the entire pipeline from stalling due to a single timeout. For scenarios sensitive to material fidelity, developers should set timeout thresholds and fallback-model switching logic at the call layer to prevent critical tasks from being interrupted by service unavailability.
Developers that depend on this model need to monitor the automatic re-run results. If the re-run still fails, they should assess whether to adjust timeout configurations in the production environment or add multi-region deployment.
Strategic Assessment
This anomaly was caused by an API failure and is unrelated to model capability degradation, so there is no need to doubt Doubao Pro's core performance. If the next Smoke evaluation re-run succeeds and scores return to normal, this incident can be confirmed as an isolated technical problem; if similar missing data persists after the re-run, stability of the call chain needs to be verified.
Based on the existing score comparison, Doubao Pro's absence from the current main leaderboard is purely a technical recording issue and does not constitute a basis for capability assessment. When selecting vendors, enterprises may continue to refer to historical complete evaluation data, while strengthening API fault tolerance during integration.
Data source: YZ Index | Run #274 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接