In the 2026-10-02 YZ Index Smoke quick test, GLM-4.6 scored 67.94 on the main leaderboard, 41.70 on code execution, and 100.00 on Material Constraints, with an integrity rating of fail (probe score 15.00).
The Sharpest Contrast in Score Structure
Material Constraints at 100.00 and code execution at 41.70 form a clear contrast. Full marks in Material Constraints indicate that the model can strictly follow the supplied material and correctly label sources in tasks requiring citations from long documents. The code execution score of 41.70 shows a low pass rate in real Python sandbox runs, below the code execution performance of multiple models that day.
An integrity rating of fail is an independent signal. 42 canary probes detected the model treating fabricated entities as genuine cited sources, and the probe score was only 15.00. That day, Claude Opus 4.7, Gemini 2.5 Pro, and Gemini 3.1 Pro all scored 90.00 on the probes; GPT-6 Sol scored 80.00; DeepSeek V4 Pro scored 80.00. GLM-4.6 was the only fail model.
Probe Trigger Mechanism Analysis
The integrity probe and the Material Constraints dimension are independent. Material Constraints examines whether the model answers based on the given material; GLM-4.6's full marks on this item indicate that it performs according to standard within controlled material. A probe fail points to the model actively fabricating external sources or fabricated entity citations in its answers, unrelated to question difficulty.
Historical records show GLM-4.6 previously triggered fail on 2026-09-29 Run#343 (probe 25.00), then turned pass on 2026-10-01 Run#354 (probe 90.00). This fail again indicates recurring integrity issues.
Specific Implications for Users
Teams that rely heavily on code execution should note GLM-4.6's 41.70 pass rate. In scenarios requiring frequent Python scripts, data processing, or algorithm testing, the model has a relatively high probability of errors.
For scenarios sensitive to fidelity to supplied material, GLM-4.6's Material Constraints score of 100.00 provides some assurance. But an integrity fail means that in tasks requiring external citations or source review, the model may insert fabricated entities or incorrect sources, increasing follow-up manual review costs.
Developers who rely on the model to generate reports or technical documents should add an extra source review step. The low probe score of 15.00 directly reflects this risk.
Strategic Assessment
GLM-4.6 achieves full marks in the Material Constraints dimension, showing that its citation ability under controlled input has reached a relatively high level. However, the code execution score of 41.70 and the integrity fail together form a clear weakness, putting it at a disadvantage in scenarios requiring high consistency and auditable output.
The other 13 models that day all received pass or warn integrity ratings; GLM-4.6 was the only fail case. This signal is worth tracking continuously in subsequent Smoke quick tests, observing fluctuations in its code execution and integrity probe scores.
Based on current data, GLM-4.6 is better suited to specific tasks with high Material Constraints requirements and low code execution needs. For teams that need both reliable code execution and trustworthy sources, the current score structure shows it is not the optimal choice.
Data source: YZ Index | Run #356 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接