GLM-4.6 scored "-" across all five dimensions—execution, grounding, judgment, integrity, and communication—in today's Smoke evaluation, with main leaderboard scores also missing. It will not participate in this period's rankings.
Data Facts
The Smoke evaluation consists of a fixed set of 10 questions daily, with 2 questions per dimension. Today, all auditable and side-leaderboard dimensions for GLM-4.6 returned no results due to API failure or timeout, and the system has triggered an automatic rerun. The comparison column for yesterday also shows "-", indicating this interruption is a single event.
Cause Analysis
The score comparison column shows all dimensions going from "-" to "-", with no numerical fluctuation whatsoever. Question randomization is within normal range for the Smoke evaluation, but the simultaneous absence of all five dimensions points to a technical interruption at the API call level rather than degradation in model response quality. The activation of the automatic rerun mechanism further confirms the cause as interface timeout or server-side response failure.
GLM-4.6's five-dimension data in today's Smoke evaluation is completely missing due to API failure/timeout. An automatic rerun has been initiated, and it will not participate in this period's rankings.
Implications for Users
Teams relying on the GLM-4.6 API for code execution tasks need to add timeout retry and fallback model switching logic in critical workflows. In scenarios sensitive to material constraints, a single API interruption can cause an entire batch of tasks to fail. It is recommended to monitor call success rates and set circuit-breaker thresholds.
- Teams heavy on code execution should hold off on high-value batch processing tasks until today's rerun results are published.
- Scenarios sensitive to material fidelity should prepare local caching or parallel calls to other available models.
Strategic Judgment
This anomaly is a clear API technical failure, not a signal of model capability degradation. After the automatic rerun completes, results can be directly compared with historical Smoke scores; today's missing data need not be interpreted as a stability issue. Only if multiple dimensions are simultaneously missing again in the next Smoke evaluation would further verification of API service reliability be warranted.
Data source: YZ Index | Run #328 | View raw data
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接