GLM-4.6 Material Constraint Score Plummets 27.3 Points, Main Score Rises 30.2 Points

In today's Smoke evaluation, GLM-4.6's material constraint score dropped from 75.00 to 47.70 points, while its main score rose from 46.29 to 76.47 points.

Score Comparison Facts

Code execution rose from 22.80 to 100.00 points, material constraint fell from 75.00 to 47.70 points, engineering judgment fell from 100.00 to 63.90 points, task expression fell from 66.70 to 50.00 points, the main score rose from 46.29 to 76.47 points, and the integrity rating shifted from pass to warn.

Dimension Composition and Fluctuation Mechanism

The main score comprises only two dimensions: code execution and material constraint. Today, material constraint declined by 27.3 points, but code execution rose by 77.2 points; after averaging the two, the main score still gained 30.2 points. The Smoke evaluation covers only 10 questions per day—2 per dimension—so day-to-day variation from random question selection falls within the normal range. The material constraint decline most likely stems from the day's randomly drawn questions placing higher demands on material fidelity, rather than an overall degradation in model capability.

The combination of a 100.00-point code execution score and a 47.70-point material constraint score directly lifted the main score to 76.47 points.

Cause Analysis

The material constraint dimension measures the model's fidelity to provided materials. The drop from 75.00 to 47.70 points indicates that, on the day's questions, the model produced more responses that deviated from the source material. The code execution score rising from 22.80 to 100.00 points shows the model performed consistently and correctly on code-related questions. Engineering judgment and task expression—the two side dimensions—also declined, by 36.1 and 16.7 points respectively, further confirming that the day's question difficulty distribution differed from yesterday's. The integrity rating shifting from pass to warn indicates that the model exhibited recordable integrity issues in some answers.

Implications for Users

Teams with heavy reliance on code execution can continue using GLM-4.6, as the code execution dimension has reached 100.00 points. For scenarios sensitive to material fidelity, caution is advised—a material constraint score of 47.70 points means the model is more likely to generate content inconsistent with the given materials. Developers who depend on engineering judgment and task expression should add manual output validation, as both side dimensions have declined noticeably.

Strategic Assessment

The main score increase was driven largely by the code execution dimension alone, while the 27.3-point drop in material constraint has partially offset that advantage. The integrity rating shifting to warn is the only signal from this evaluation that warrants continued monitoring. If material constraint remains below 60 points in the next Smoke evaluation, genuine model degradation should be considered; if it rebounds above 70 points, this fluctuation is more likely attributable to random question selection.

This evaluation reflects only a single day's performance across 10 questions and cannot support a conclusion of sustained degradation. Enterprises selecting models should reassess GLM-4.6's overall stability only after the material constraint dimension recovers to above 75 points.


Data source: YZ Index | Run #256 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!