GLM-4.6 Probe Score 15.00, Integrity Fail, Main Leaderboard 48.25, Material Constraint 46.10

In Run #307 on 2026-09-03, GLM-4.6 received an integrity rating of fail, with a probe score of 15.00, a main leaderboard score of 48.25, code execution at 50.00, and material constraint at 46.10.

Score Structure Breakdown

This Smoke quick test's main leaderboard covers only two dimensions: code execution and material constraint. GLM-4.6 scored 50.00 on code execution and 46.10 on material constraint, both below 60, jointly pulling the main leaderboard score down to 48.25. The integrity probe score of 15.00 is an independent signal and has no direct correlation with the two dimensions above.

All other models tested the same day passed the integrity rating: GPT-o3 probe 80.00, Doubao Pro 80.00, GPT-5.5 90.00, Claude Sonnet 4.6 90.00, Gemini 3.1 Pro 80.00, Grok 4 80.00, Claude Opus 4.7 100.00, Qwen3 Max 90.00, Gemini 2.5 Pro 90.00, DeepSeek V4 Pro 80.00.

Mechanism Behind the Integrity Fail

The integrity rating uses 42 canary probes to detect whether a model treats fabricated entities as real citations. GLM-4.6's probe score of 15.00 indicates that it fabricated sources or citations on at least some of the probes. The Smoke test explicitly states that probe triggering is an independent signal unrelated to question difficulty.

Historical data shows that GLM-4.6 also failed in Run #304 on 2026-09-01 (probe 25.00) and received a warn in Run #305 on 2026-09-02 (probe 45.00). Across the three tests, it triggered a fail twice, with probe scores remaining persistently low.

Specific Impact on Users

A material constraint score of 46.10 indicates that the model failed to strictly answer based on the given materials and cite correctly in long-document citation verification tasks. Enterprise document processing, legal contract review, and academic literature organization scenarios that rely on precise citation will face higher risks.

A code execution score of 50.00 indicates a somewhat below-average pass rate in the real Python sandbox. Developers who need to reliably run data analysis scripts or automated tests will need to add extra manual verification steps.

The integrity fail directly affects scenarios that require verifiable sources. If outputs such as financial research reports, policy interpretations, and product white papers contain fabricated citations, this will lead to compliance and trust issues.

Strategic Assessment

The combination of GLM-4.6's current main leaderboard score of 48.25 and material constraint score of 46.10 reflects that both its performance under strict material constraints and its code execution capability fall below the passing threshold. The independent failure signal of the 15.00 probe score points to a systemic deficiency in the model's control over citation authenticity.

Models that passed the integrity rating on the same day all scored 80 or above on probes, forming a stark contrast. GLM-4.6's repeated fail records warrant continued observation to determine whether this is sporadic fluctuation or a persistent issue.

For enterprises currently in the model selection process, GLM-4.6 is suitable for internal draft generation scenarios with lower citation accuracy requirements. For production environments sensitive to material fidelity, it is recommended to prioritize models with probe scores above 90 and retain manual verification workflows.


Data source: YZ Index | Run #307 | View Raw Data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!