GLM-4.6 scored 74.00 on the main leaderboard, 82.30 on code execution, 95.00 on material constraint, and 70.80 on engineering judgment in today's Smoke evaluation. The task expression dimension has no data, and the completeness and communication dimensions are missing due to API failure/timeout. These have entered automatic retesting and do not participate in this period's ranking.
Data Facts: Single-Day Scores and Dimension Completeness
The Smoke evaluation features only 10 questions per day, with 2 questions per dimension. Today, GLM-4.6 scored 82.30 on code execution, 95.00 on material constraint, and 70.80 on engineering judgment, with a main leaderboard score of 74.00. The task expression dimension is vacant. All comparison values against yesterday are "-", indicating this is the first time these figures have been recorded. The 13.7-point gap between material constraint (95.00) and code execution (82.30), while engineering judgment (70.80) trails both.
Cause Analysis: Distinguishing API Failure from Question Draw Fluctuation
The score comparison shows that the task expression dimension is directly missing, and integrity and communication also failed to appear, with the cause officially attributed to API failure/timeout. This differs from normal single-day fluctuations caused by question draw randomness. Question draw fluctuation typically affects the score levels of specific questions but does not result in complete data absence across an entire dimension. The 13.2-point gap between engineering judgment (70.80) and material constraint (95.00) reflects performance variation across dimensions under a limited question set, but the missing dimensions themselves point to API-level technical issues rather than simultaneous degradation of model capabilities across all dimensions.
The API failure directly prevented data collection for two dimensions, and the main leaderboard score of 74.00 was calculated solely from the returned code execution and material constraint results.
Implications for Users
For code-focused teams calling GLM-4.6, the 82.30 score indicates the model remains usable for code-related tasks, but API timeout retry mechanisms should be prepared. In scenarios sensitive to material fidelity, the 95.00 score provides strong reference value and warrants priority consideration for short-text constraint tasks. Developers relying on engineering judgment should note that the 70.80 score may introduce inconsistency risks, particularly in scenarios requiring multi-round verification.
- When the API is unstable, missing dimensions directly undermine the reliability of model selection decisions.
- The automatic retest mechanism can mitigate a single failure, but developers still need to monitor consecutive call success rates.
Strategic Assessment: Technical Failure Rather Than Capability Degradation
Based on the current score comparison, GLM-4.6's anomaly this period is primarily attributable to API failure rather than genuine model degradation. The material constraint score of 95.00 and code execution score of 82.30 remain within a usable range, while the engineering judgment score of 70.80 exposes volatility under a limited sample. The next period will need to verify whether API call success rates recover and whether full-dimension scores return to normal ranges after retesting. If API issues persist, the model's viability in production environments will be directly constrained.
Single-day fluctuation is normal in Smoke evaluations, but missing dimension data is an observable technical signal. GLM-4.6's main leaderboard score of 74.00 this period should be treated as reference only; a complete ranking awaits the return of retested data.
Data source: YZ Index (赢政指数) | Run #257 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接