In Run #307 on 2026-09-03, GLM-4.6 received an integrity rating of fail, with a probe score of 15.00, a main leaderboard score of 48.25, code execution at 50.00, and material constraint at 46.10.
Score Structure Breakdown
This Smoke quick test's main leaderboard covers only two dimensions: code execution and material constraint. GLM-4.6 scored 50.00 on code execution and 46.10 on material constraint, both below 60, jointly pulling the main leaderboard score down to 48.25. The integrity probe score of 15.00 is an independent signal and has no direct correlation with the two dimensions above.
All other models tested the same day passed the integrity rating: GPT-o3 probe 80.00, Doubao Pro 80.00, GPT-5.5 90.00, Claude Sonnet 4.6 90.00, Gemini 3.1 Pro 80.00, Grok 4 80.00, Claude Opus 4.7 100.00, Qwen3 Max 90.00, Gemini 2.5 Pro 90.00, DeepSeek V4 Pro 80.00.
Mechanism Behind the Integrity Fail
The integrity rating uses 42 canary probes to detect whether a model treats fabricated entities as real citations. GLM-4.6's probe score of 15.00 indicates that it fabricated sources or citations on at least some of the probes. The Smoke test explicitly states that probe triggering is an independent signal unrelated to question difficulty.
Historical data shows that GLM-4.6 also failed in Run #304 on 2026-09-01 (probe 25.00) and received a warn in Run #305 on 2026-09-02 (probe 45.00). Across the three tests, it triggered a fail twice, with probe scores remaining persistently low.
Specific Impact on Users
A material constraint score of 46.10 indicates that the model failed to strictly answer based on the given materials and cite correctly in long-document citation verification tasks. Enterprise document processing, legal contract review, and academic literature organization scenarios that rely on precise citation will face higher risks.
A code execution score of 50.00 indicates a somewhat below-average pass rate in the real Python sandbox. Developers who need to reliably run data analysis scripts or automated tests will need to add extra manual verification steps.
The integrity fail directly affects scenarios that require verifiable sources. If outputs such as financial research reports, policy interpretations, and product white papers contain fabricated citations, this will lead to compliance and trust issues.
Strategic Assessment
The combination of GLM-4.6's current main leaderboard score of 48.25 and material constraint score of 46.10 reflects that both its performance under strict material constraints and its code execution capability fall below the passing threshold. The independent failure signal of the 15.00 probe score points to a systemic deficiency in the model's control over citation authenticity.
Models that passed the integrity rating on the same day all scored 80 or above on probes, forming a stark contrast. GLM-4.6's repeated fail records warrant continued observation to determine whether this is sporadic fluctuation or a persistent issue.
For enterprises currently in the model selection process, GLM-4.6 is suitable for internal draft generation scenarios with lower citation accuracy requirements. For production environments sensitive to material fidelity, it is recommended to prioritize models with probe scores above 90 and retain manual verification workflows.
Data source: YZ Index | Run #307 | View Raw Data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接