Doubao Pro experienced complete data loss across all five dimensions — execution, grounding, judgment, integrity, and communication — in today's Smoke evaluation, with API failure or timeout as the direct cause. It will not participate in the main leaderboard ranking this cycle.
Score Comparison Shows Completely Blank
The score comparison from yesterday to today is entirely "-"; code execution, material grounding, engineering judgment, task expression, and main leaderboard all show no specific values. This indicates that the evaluation process was interrupted at the data collection stage, not that the model received low scores after answering questions.
Possible Cause Analysis
Question-pool sampling fluctuations typically manifest as minor variations in a single dimension, whereas this time all dimensions simultaneously lack data, pointing to an API connection or response timeout. Genuine model degradation would be reflected in declining scores on specific questions; no answering evidence was recorded this time, so a degradation conclusion cannot be supported.
Data loss caused by API failure is unrelated to the model's intrinsic capability; the automatic re-run mechanism has been initiated.
Implications for Users
Teams that rely heavily on code execution should add timeout retry logic when calling Doubao Pro, to avoid a single API call failure disrupting business processes. For scenarios sensitive to material faithfulness, it is advisable to wait for re-run results before formal deployment to confirm whether the grounding dimension has returned to normal.
- When selecting vendors, enterprises may place Doubao Pro on a watch list and make a final judgment after the re-run data is published.
- Developers depending on the model's continuous service should prepare a backup model to handle occasional API timeouts.
Strategic Assessment
This anomaly is a technical-level signal and does not constitute evidence of model capability degradation. If the next re-run data restores complete dimensional records, no further attention is needed; if the data is missing again, it will be necessary to verify whether API stability has become a persistent bottleneck.
The stability dimension measures the consistency of the model's answers, calculated based on the standard deviation of scores. Since no scores were available this time to compute a standard deviation, no stability conclusion can be drawn.
Data discipline requires that all judgments be based solely on provided comparisons; this cycle's comparison is blank, so conclusions are limited to the technical failure itself. Users can continue to track the automatic re-run results to obtain auditable execution and grounding values.
Engineering judgment and task expression, as side-leaderboard dimensions, are also missing this time and likewise do not affect the main leaderboard ranking logic. The integrity rating "pass" remains the entry threshold, and this incident did not touch that threshold.
Availability, as an operational signal, indicates transient unavailability on the API side this time; it is recommended to set up health checks in production environments.
In summary, the anomaly in Doubao Pro's Smoke evaluation this cycle was caused by an API failure, and the model's true capability shows no signs of degradation. Waiting for re-run data is the only reasonable action based on current records.
Data source: YZ Index | Run #264 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接