GLM-4.6 Integrity Rating Drops from Pass to Fail, Code Execution Surges by 47 Points

In today's Smoke evaluation, GLM-4.6's integrity rating directly dropped from pass to fail, while its code execution score rose from 50.00 yesterday to 97.00. The overall main ranking increased from 62.83 to 74.00.

Score Change Breakdown

Material faithfulness dropped from 78.50 to 75.00, engineering judgment fell from 60.40 to 0.00, and task expression remained unchanged at 50.00. The core main ranking is weighted only by code execution and material faithfulness, so the single-day increase of 47 points in code execution directly boosted the main ranking by 11.2 points.

Possible Cause Analysis

The Smoke evaluation has only 10 questions daily, 2 per dimension, so single-day sampling fluctuation can cause significant score swings. The 47-point leap in code execution from 50.00 is most likely due to today's two programming questions that the model excels at. The direct drop to zero in engineering judgment indicates that the model completely failed to provide valid judgments on the two side-ranking questions drawn today. Material faithfulness only slightly decreased by 3.5 points, and task expression remained unchanged, suggesting no systematic degradation in core auditable dimensions.

The integrity rating changing from pass to fail is a threshold event, not a bonus item. This change is most likely directly related to the engineering judgment score of 0.00. The model may have exhibited obvious fact fabrication or logical breaks on the side-ranking questions, causing the evaluation system to deem its integrity below standard.

Implications for Users

Development teams focused on code execution can continue using GLM-4.6 today, as its 97.00 performance is better than yesterday, but they should also enable manual code review processes. Scenarios relying on material faithfulness should be cautious: the 75.00 material faithfulness score is already lower than yesterday, and the fail integrity rating means the model output may contain unverifiable content.

Automated workflows that heavily rely on engineering judgment (e.g., requirements review, architecture decision support) should immediately switch to other models, as the 0.00 score indicates GLM-4.6 is completely ineffective on such questions.

Strategic Assessment

Based on current score comparisons, GLM-4.6's main ranking increase is driven primarily by the single dimension of code execution, not an overall capability improvement. The simultaneous occurrence of a fail integrity rating and zero engineering judgment suggests a higher probability of single-day sampling fluctuation rather than genuine model degradation. The next Smoke evaluation should focus on verifying whether engineering judgment and integrity rating recover, to confirm the nature of the fluctuation.

Current data does not support judging GLM-4.6 as persistently overvalued or undervalued. The only clear signal is that its extreme instability on side-ranking dimensions deserves ongoing monitoring.


Data source: YZ Index | Run #241 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!