Doubao Pro Code Execution Crashes 80 Points, Main Score Drops 41.2 in a Single Day

Doubao Pro Code Execution Crashes 80 Points, Main Score Drops 41.2 in a Single Day

Doubao Pro in today's Smoke evaluation saw its main score directly drop from 81.33 to 40.12, a decline of 41.2 points. The core reason is that the code execution dimension collapsed from a perfect score of 100 to 20, losing 80 points in a single day.

Sampling Fluctuation or True Degradation

Smoke evaluation only has 2 questions per day, and an extremely low score in the code execution dimension typically indicates that the sampling concentrated on high-difficulty or edge-case scenarios. The material constraint dimension actually rose from 58.5 to 64.7, suggesting no systemic degradation in constraint following. The engineering judgment score dropped from 38.4 to 10, also pointing to the same batch of questions potentially biased toward complex multi-step reasoning.

But a single-dimension drop of 80 points has already exceeded the normal sampling range. The YZ Index stability dimension shows that Doubao Pro currently scores only 31.7, meaning the standard deviation of scores on similar historical questions is extremely large and consistency is low. This crash is more likely an unstable performance of the model in specific code scenarios rather than a collapse of overall capability.

Comparison of Recent Industry Trends

ByteDance has recently invested Doubao's main resources in cost-performance optimization and Chinese long-text scenarios, and code ability has not become a key iteration direction. Meanwhile, open-source models such as DeepSeek-Coder-V2 and Qwen2.5-Coder have continuously released targeted updates, pushing Doubao's relative position on pure code tasks backward. The extremely low code execution score in today's test is consistent with its product strategy focus.

The trust rating changed from warn to pass, indicating that the model did not show obvious hallucinations or violations in this response, and basic reliability remains at the passing line.

Need for Continued Attention

A single-day drop of 41.2 points requires tracking the data of the next 3 days. If the code execution dimension is below 40 for two consecutive days, it can be judged as true capability fluctuation; if it rebounds quickly tomorrow, it can be basically attributed to question sampling. It is currently recommended to lower the priority of code-related tasks for Doubao Pro and wait for stability data to stabilize.

The 80-point single-dimension crash does not expose a model collapse, but rather the severe fluctuations that have long existed in its code scenarios.

Data source: YZ Index | Run #136 | View raw data