GPT-o3 Code Execution Plunges 24.7 Points, Main Leaderboard Falls to 78.13 — Smoke Evaluation Anomaly Warrants Attention

In today's Smoke evaluation, GPT-o3's code execution score fell from yesterday's 94.50 to 69.80, a 24.7-point drop, and the main leaderboard overall declined from 86.04 to 78.13, down 7.9 points. The anomaly is most likely due to question-sampling variance rather than genuine model degradation, though code-intensive users should stay vigilant.

GPT-o3 Code Execution Smoke Test
223

Doubao Pro's Material Constraint Score Plunges 27.6 Points While Code Execution Soars 49.3 Points

In today's Smoke evaluation, Doubao Pro's material constraint score plunged 27.6 points to 58.30, while its code execution score soared 49.3 points to 99.30, lifting its main leaderboard score from 66.16 to 80.85. The analysis attributes these dramatic opposing swings primarily to question sampling fluctuation rather than genuine model degradation.

Doubao Pro Material Constraints Smoke Test
270