GLM-4.6: Material Constraint +28.2 but Engineering Judgment −39.6; Integrity Upgraded from Fail to Pass, Main Leaderboard Rises 12.7

In today's Smoke evaluation, GLM-4.6 held its code execution score steady at 50.00, while material constraint rose from 46.10 to 74.30. As a result, the main leaderboard score climbed from 48.25 to 60.94, and the integrity rating improved from fail to pass.

Data Breakdown

Comparing yesterday's and today's scores, code execution stayed completely flat at 50.00, while material constraint gained 28.2 points in a single day. Engineering judgment fell from 64.60 to 25.00, and task expression fell from 50.00 to 25.00. The main leaderboard rose 12.7 points overall, driven primarily by the improvement in material constraint.

Cause Analysis

The Smoke evaluation uses only 10 questions per day — 2 questions per dimension — so the sample size is extremely small, and variance in the random question draw is the most direct mechanism behind score changes. The sharp rise in material constraint indicates that the model performed better on the material-fidelity requirements of the two questions drawn today. At the same time, engineering judgment fell 39.6 points and task expression fell 25 points on the secondary dimensions, suggesting the drawn questions may have placed higher demands on engineering decision-making and task specification, causing the model to lose more points on those items. The integrity rating shifting from fail to pass reflects that the model triggered no integrity deductions on today's drawn questions.

Material constraint at 74.30 versus engineering judgment at 25.00 — a 38.3-point gap between dimensions for the same model on the same day — points to the question draw rather than to model architectural regression.

Implications for Users

Developers in material-constraint-sensitive scenarios may consider running tasks that require strict fidelity to source texts on today's top-scoring model. Teams that rely heavily on engineering judgment should note that the 25.00 secondary score may pose a risk of inconsistent decision-making, and human review at critical engineering steps is recommended. Scenarios dependent on task expression face the same 25.00-level score, so outputs may be less structured.

Strategic Assessment

The improvement to a 60.94 main leaderboard score was driven entirely by the material constraint dimension, while code execution held at 50.00, indicating no systematic change in the model's core execution capability. The simultaneous sharp declines in engineering judgment and task expression are typical results of a small-sample question draw, and there is currently no evidence of genuine model regression. The integrity rating moving to pass is a positive signal, but the severe volatility on the secondary dimensions means the next Smoke evaluation should be watched to see whether engineering judgment recovers to above 60.

This evaluation again confirms that under Smoke's 10-question daily design, standard deviations on the secondary dimensions can easily exceed 30 points, while the main leaderboard remains relatively stable. Enterprises selecting models should base their decisions on the three-day average of the main leaderboard rather than on any single day's data.


Data source: YZ Index | Run #308 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!