GLM-4.6 Smoke Test: Material Constraint Scores 71.90, Code Execution and Integrity Dimensions Missing

GLM-4.6 scored 71.90 on material constraint, 63.90 on engineering judgment, and 50.00 on task expression in today's Smoke evaluation. Both the code execution and integrity dimensions were entirely absent due to API failures or timeouts, and the model is not ranked on the main leaderboard.

Data Facts

Compared with yesterday, only three side leaderboards or partial dimension scores were retained today, with the execution dimension and integrity rating left completely empty. The Smoke evaluation runs a fixed 10 questions per day, with 2 questions per dimension; a large single-day standard deviation is a known characteristic, but this absence is not a score decline—rather, entire dimensions of data were not returned.

Cause Analysis

The questions corresponding to the missing execution and integrity dimensions typically require the model to complete code snippets or strictly follow material instructions. API timeouts are most likely to occur on requests that require multi-round interaction or longer contexts. The fact that material constraint (71.90) and engineering judgment (63.90) were still returned indicates that the model's basic generation capability was not fully interrupted—the issue lies in interface stability rather than degradation at the parameter level. The low task expression score of 50.00 may be related to the question draw containing more open-ended expression requirements, but without execution dimension data, it is impossible to determine whether this is a systematic decline.

Dimension absence caused by API failures and true model capability degradation are two different mechanisms; the former can be recovered through re-runs, while the latter requires multiple days of continuous data to confirm.

Implications for Users

Enterprises that heavily depend on code execution scenarios should immediately inspect GLM-4.6 API call logs, focusing on timeout and retry rates. If daily workflows include material fidelity requirements, 71.90 can still serve as a reference baseline, but the missing execution dimension means the model's actual performance on code generation tasks cannot be verified. Developers should hold off on using GLM-4.6 as a core code execution node until the next round of re-run results are available.

Strategic Assessment

This anomaly was primarily caused by the API layer rather than a collapse in model capability. It is worth continuing to monitor interface stability rather than jumping to conclusions about model degradation. It is recommended that the next Smoke evaluation focus on verifying whether the execution dimension has recovered and whether the integrity rating has reappeared; if the execution dimension remains persistently missing after re-runs, further investigation into API rate limiting or server-side configuration issues is warranted.

The inherent volatility of a 10-question single-day test is normal; this incident is more of an operational signal than a capability signal. When selecting models, enterprises should treat API availability as an independent consideration, evaluated alongside main leaderboard scores.


Data source: YZ Index | Run #264 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!