GLM-4.6 Code Execution Drops 12.5 Points, Perfect Material Constraint Score Lifts Main Leaderboard by 5.4

GLM-4.6 scored 79.38 points on the main leaderboard in today's Smoke evaluation, with 62.50 points in the code execution dimension and 100.00 points in the material constraint dimension. Its integrity rating changed from fail to pass.

Score Comparison Facts

Yesterday's scores were 75.00 for code execution, 78.30 for material constraints, 75.00 for engineering judgment, and 20.00 for task expression, with a main leaderboard score of 74.00. Today's scores are 62.50 for code execution, 100.00 for material constraints, 75.00 for engineering judgment, and 41.70 for task expression, with a main leaderboard score of 79.38. Code execution fell 12.5 points in a single day, material constraints rose 21.7 points, and task expression rose 21.7 points.

Analysis of Fluctuation Causes

The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension, so the randomness of question selection directly affects scores. The drop in code execution from 75.00 to 62.50 is most likely due to today's draw containing questions that demand higher code correctness, while the material constraint dimension drew questions that are easier to satisfy. Engineering judgment remained unchanged at 75.00 points, indicating that question difficulty in this dimension is relatively stable. The integrity rating changing from fail to pass likewise points to daily variation in question selection rather than systematic issues with model parameters or training data.

If a model were truly degrading, one would typically observe sustained declines across multiple dimensions simultaneously. Today, while code execution declined, both material constraints and task expression rose significantly, and the overall main leaderboard score improved by 5.4 points. This pattern is more consistent with random question selection than with model capability degradation.

Specific Implications for Users

Teams that rely heavily on code execution should build in a verification step to account for single-answer fluctuations when adopting GLM-4.6. In material constraint scenarios, today's perfect 100.00-point performance demonstrates the model's advantage in strictly adhering to input materials, making it suitable for document processing tasks requiring high fidelity. Task expression rose from 20.00 to 41.70, indicating improved performance under today's question draw in scenarios requiring clear output structures.

Developers relying on GLM-4.6 for code generation should increase unit test coverage to mitigate potential day-to-day fluctuations in the code execution dimension. For enterprise selection, the integrity rating moving to pass means the model has cleared the basic integrity threshold and can proceed to the next stage of evaluation.

Strategic Assessment

Based on today's data, the inverse movement of GLM-4.6's code execution score (62.50) and material constraint score (100.00) is most likely caused by question selection fluctuations rather than genuine model degradation. The main leaderboard score of 79.38, up from yesterday, indicates no systematic decline in overall capability. The next evaluation round should monitor whether code execution recovers to around 75.00; only if it remains at a low level for two consecutive days should further verification be considered.

There is currently no need to launch additional stability testing for GLM-4.6; routine monitoring is sufficient. Both the engineering judgment score holding at 75.00 and the integrity rating changing to pass do not support the conclusion of model capability decline.


Data source: YZ Index | Run #275 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!