In today's Smoke evaluation, GLM-4.6's Material Constraint score fell from 74.30 to 53.30, a decline of 21 points; its Code Execution score rose from 50.00 to 75.00; and its main leaderboard score moved up from 60.94 to 65.24.
Score Comparison and Direct Facts
The day-over-day changes from yesterday to today are as follows: Code Execution 50.00→75.00, Material Constraint 74.30→53.30, Engineering Judgment 25.00→95.80, Task Expression 25.00→50.00, Main Leaderboard 60.94→65.24, Integrity Rating pass→warn. The Smoke evaluation covers only 10 questions per day, with 2 questions per dimension, so single-day score fluctuations fall within the normal range.
Likely Causes
The 21-point plunge in the Material Constraint dimension most likely stems from a mismatch between the difficulty or type of the two Material Constraint questions drawn that day and the model's current output style. The simultaneous 25-point rise in Code Execution suggests the model responded more accurately to task requirements in that dimension. The dramatic swing in Engineering Judgment from 25.00 to 95.80 further supports that question-draw randomness is the primary driver, rather than any overall degradation of model capability.
The Integrity Rating shifting from pass to warn indicates that this run may have produced responses inconsistent with the source material or containing fabricated details, which is directly tied to the decline in Material Constraint. Task Expression rose 25 points, showing improvement in the clarity of task descriptions, but the gain was not enough to offset the loss in Material Constraint.
Implications for Users
For enterprise scenarios that rely heavily on fidelity to source material—such as contract review, policy interpretation, or technical documentation generation—GLM-4.6's performance today signals a need for additional human verification. The improved Code Execution score, by contrast, offers a positive signal for developers who need to generate runnable code quickly.
For teams that require the model to assess solution feasibility, the sharp rise in Engineering Judgment suggests today's data may understate its actual capability. The main leaderboard rose only 4.3 points, indicating that the loss in Material Constraint was partially offset by the gain in Code Execution. Still, the warn Integrity Rating already constitutes a clear caution for model selection.
Strategic Assessment
Based on the current score comparison, the simultaneous 21-point drop in Material Constraint and 70.8-point jump in Engineering Judgment is most plausibly explained by question-draw fluctuation rather than genuine model regression. With the Smoke evaluation's small daily sample, a single-day anomaly is insufficient to establish a directional change in model capability.
We recommend that the next evaluation cycle focus on whether Material Constraint recovers to above 70 and whether the Integrity Rating returns to pass. If Material Constraint stays below 55 for two consecutive days while the Integrity Rating remains at warn, GLM-4.6 should be deprioritized for material-sensitive scenarios.
Current data does not support a long-term negative verdict on the model. However, the warn Integrity Rating already constitutes a short-term usage risk signal; developers relying on this model should add a secondary verification step for tasks involving material constraints.
Data source: YZ Index | Run #309 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接