Grok 4 Main Score Plunges 8.4 Points, Material Constraint Drops 17.6 Points in a Single Day
Grok 4's main score in today's Smoke evaluation dropped 8.4 points from 87.66 to 79.30, with the Material Constraint dimension falling 17.6 points.
Grok 4's main score in today's Smoke evaluation dropped 8.4 points from 87.66 to 79.30, with the Material Constraint dimension falling 17.6 points.
On 2026-07-10, the YZ Index Smoke Quick Test covered 9 models, with GPT-o3 ranking first at 86.9 points. Smoke is a daily 10-question quick test for monitoring short-term signals, not equivalent to Full weekly rankings.
GPT-o3's material constraint score dropped 16.8 points in today's Smoke evaluation, while task expression fell 28.3 points, causing the main ranking total to decline from 83.44 to 80.39.
In today's Smoke evaluation, Qwen3 Max's Material Constraint score dropped from 83.60 to 68.50, a decrease of 15.1 points, while its Code Execution score rose from 73.10 to 91.50, and its main ranking score increased from 77.83 to 81.15.
On July 9, 2026, the YZ Index Smoke Quick Test covered 10 models, with Claude Opus 4.7 ranking first with a score of 90.51. The Smoke test is a daily 10-question quick assessment for monitoring short-term signals, not equivalent to the Full weekly ranking.
Grok 4 WDCD scores 95.00, up 3.8 points from Run #211, maintaining first place; DeepSeek V4 Pro jumps 26.2 points to 94.00, and GLM-4.6 rises 21.8 points to 93.60, both within 2 points of Grok 4. Only Claude Sonnet 4.6 declines, by 5.9 points.
In the WDCD v3.1 pilot, the Business Rules scenario scored the lowest overall, with champion claude-opus-4.7 achieving only 3.5/4 and bottom-ranked qwen3-max scoring just 1.3/4, far below the champion scores of the other four scenarios.
In a worst-of-3 sampling of only 8 v2 anchor questions, the average R3 integrity rate across 11 models was merely 61.4%, while R1 confirmation rate remained as high as 95% and R2 resistance rate 73%, revealing the true performance of mainstream models under hard constraints.
Grok 4 leads the WDCD Compliance Leaderboard with 95.00 points, while Claude Sonnet 4.6 ranks 11th with 64.10 points, a gap of 30.9 points.
The 2026-07-08 YZ Index Smoke Quick Test covered 10 models, with DeepSeek V4 Pro ranking first at 95.19 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.
On 2026-07-07, the Winzheng YZ Index Smoke Quick Test covered 11 models. Claude Opus 4.7 and Grok 4 tied for first place with a score of 96.99.
In the YZ Index Smoke Quick Test on July 6, 2026, Doubao Pro ranked first with a Main Board score of 83.91, covering 11 models in 10 daily questions. The test focuses on code execution and material constraints, serving as a short-term monitoring signal rather than a long-term conclusion.
In the Smoke Quick Test Run#214 on 2026-07-05, GLM-4.6 scored 60.04 on the main leaderboard, with code execution at 88.70, material constraint at 25.00, integrity rating fail, and probe score 0.00.
On July 5, 2026, the YZ Index Smoke Quick Test covered 11 models, with Doubao Pro and Gemini 3.1 Pro tying for first place at 88.54 points. Smoke is a daily 10-question quick test for observing short-term signals and is not equivalent to the Full weekly ranking.
Agent-Assisted SGLang Development: An Initial ExplorationSGLang TeamJuly 2, 2026SGLang development increasingly goes beyond isolated code changes. The same repository now spans LLM serving,…
In the June 2026 Smoke evaluation of the YZ Index, Qwen3 Max's main leaderboard score fell from 84.92 to 72.02, a drop of 12.9 points, with the code execution dimension plummeting from 96.30 to 69.50.
In the Smoke Lite evaluation of 11 models on July 4, 2026, by the YZ Index, Gemini 2.5 Pro ranked first with a Main Board score of 96.99, while Qwen3 Max's Main Board score plunged 12.9 points to 72.02.
In the WDCD v3.1 compliance test, the business rules scenario scored the lowest among all models, with grok-4 leading at 3.5/4, while doubao-pro and qwen3-max only scored 1.55/4.
In 275 samples on 8 v2 anchor questions, the average R1 confirmation rate was 0.99, but the R3 integrity rate was only 30.2%, with 44 complete collapses (score 0). This data directly reveals the rapid degradation pattern of models after initial commitment as rounds increase.
Grok 4 tops the WDCD Compliance Leaderboard with 91.20 points, while Qwen3 Max ranks last with 57.48 points, a gap of 33.72 points between the top and bottom.