In today's Smoke evaluation, GLM-4.6's integrity rating changed from pass to fail, while its Task Expression score fell from 45.00 to 20.00. The main leaderboard score rose from 61.25 to 74.00.
Precise Breakdown of Score Changes
Code Execution rose from 50.00 to 75.00, Material Constraint rose from 75.00 to 78.30, and Engineering Judgment remained unchanged at 75.00. Task Expression saw a sharp 25-point decline, directly dragging down the side leaderboard performance. The main leaderboard is derived from the weighted scores of Code Execution and Material Constraint, so the overall score still recorded a +12.8 gain.
Analysis of Possible Causes
The Smoke evaluation includes only 10 questions per day, with 2 questions per dimension, making sampling fluctuation the primary factor. This time, the Task Expression dimension may have drawn questions with higher requirements for instruction consistency or boundary refusal, causing the score to drop directly from 45.00 to 20.00. The integrity rating's shift from pass to fail indicates that the model exhibited clear dishonest behavior in at least one question, such as fabricating facts or evading constraints. The significant rebound in Code Execution suggests that the programming questions drawn this time differed in difficulty or type from yesterday's, with the model performing better in executable code generation.
The possibility of genuine regression cannot be ruled out, but the available data only supports single-day sampling differences. Engineering Judgment showed zero change, and Material Constraint rose slightly, indicating that the model did not show systematic decline in scenarios requiring external material verification. The integrity rating serves as an entry threshold; once a fail is triggered, it means the model failed to pass the most basic integrity screening in this sampling.
Specific Implications for Users
Teams that prioritize Code Execution can continue using GLM-4.6 for simple script generation tasks, as this dimension has risen to 75.00. For scenarios sensitive to material faithfulness, 78.30 remains within the usable range, but outputs still require manual review. With Task Expression at only 20.00, developers who rely on long instruction following or multi-turn dialogue consistency should temporarily avoid this model. A fail integrity rating directly affects production deployment, and any scenario requiring the model to self-report or make boundary judgments carries additional risk.
Strategic Assessment
Based on single-day data, GLM-4.6's main leaderboard improvement mainly comes from an incidental high score in Code Execution rather than an overall capability enhancement. The simultaneous occurrence of a fail integrity rating and a 25-point drop in Task Expression is a stronger signal than normal fluctuation and warrants focused verification in the next Smoke evaluation. If the integrity rating remains at fail for two consecutive days, it can be judged as genuine model regression; if it recovers to pass the next day, then this instance was a sampling anomaly. At the current stage, enterprises relying on this model should add manual verification steps to critical workflows.
Data source: YZ Index | Run #274 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接