In today's Smoke evaluation, GLM-4.6 scored 58.30 in code execution, 100.00 in material constraints, and 75.00 in engineering judgment, with a main leaderboard score of 74.00. Due to API failure/timeout, the integrity and communication dimensions are missing; it has entered automatic rerun and is not included in the ranking for this period.
Data Facts: Dimension Score Distribution Under 10 Questions in a Single Day
The Smoke evaluation has a fixed 2 questions per dimension each day, for 10 questions total. In this period, GLM-4.6 has scores in only three dimensions: 58.30 in code execution, 100.00 in material constraints, and 75.00 in engineering judgment. The task expression dimension is completely missing due to API timeout. The main leaderboard score of 74.00 is weighted from code execution and material constraints. All the above data comes from a single run today.
Cause Analysis: Distinguishing API Failure from Question Sampling Fluctuation
The code execution score of 58.30 contrasts sharply with the material constraints score of 100.00. A perfect material constraints score indicates the model performs stably when strictly following provided materials, and the difficulty or constraint strength of the questions themselves did not exceed the model's capabilities. The code execution score of 58.30 is notably low. Combined with the officially flagged "API failure/timeout," the most likely cause is an API connection interruption that directly caused failure in code generation or execution, rather than a degradation in the model's own understanding of code logic.
The engineering judgment score of 75.00 is in the medium range and was likewise affected by API instability, but it did not fall to the same low level as code execution. Question sampling fluctuation is normal in a daily 10-question test, but this run also recorded API timeouts and caused two dimensions of data to be missing, pointing to an external service-layer problem rather than random changes in question difficulty.
Implications for Users
Developers who rely heavily on code execution scenarios need to immediately assess GLM-4.6's performance in this period's Smoke test. A score of 58.30 means that in tasks requiring runnable code generation or multi-step debugging, the success rate may be notably lower than that of other models. The material constraints score of 100.00 shows that in scenarios with strict RAG or high document-faithfulness requirements, the model can still provide reliable output.
For enterprises selecting models, API stability has become the primary risk point in using GLM-4.6 at present. The engineering judgment score of 75.00 suggests that in scenarios requiring trade-offs among technical solutions, the model can still give medium-level recommendations, but overall reliability is weakened by the API issue.
Strategic Judgment
Based on today's scores and the API failure record, GLM-4.6's code execution score of 58.30 is most likely caused by a service-layer timeout, rather than a real degradation in model capability. The perfect material constraints score proves that it still maintains a high standard in the constraint-following dimension. It is recommended to focus on verifying the complete data after the automatic rerun next period; if API problems persist, consider switching code-execution-intensive tasks to other available models.
This anomaly has been clearly recorded as an API failure, so it does not constitute a regular sample for model capability evaluation. Enterprise users should treat GLM-4.6's current main leaderboard score of 74.00 as a temporary value affected by external factors, and make a final judgment after the rerun is completed.
Data source: YZ Index | Run #368 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接