Grok 4 Tops WDCD Commitment Ranking with 97.5 Points; Doubao Pro Trails at 68 Points
In the WDCD v3.1 commitment test, Grok 4 ranked first with 97.50 points while Doubao Pro ranked last with 68.00 points, a 29.5-point gap between the top and bottom.
In the WDCD v3.1 commitment test, Grok 4 ranked first with 97.50 points while Doubao Pro ranked last with 68.00 points, a 29.5-point gap between the top and bottom.
On 2026-08-05, the YZ Index Smoke quick test covered 9 models, with DeepSeek V4 Pro, GPT-5.5, and GPT-o3 tying for first place at 80.52 points. The briefing highlights notable single-day declines for Qwen3 Max, Grok 4, and Gemini 3.1 Pro, pending verification in subsequent runs.
GLM-4.6's Smoke evaluation today recorded 51.80 for material constraint and 50.00 for task expression, while data for the execution and judgment dimensions is missing due to an API failure. An automatic re-run has been initiated, and this round will not be included in the main leaderboard ranking.
Doubao Pro's Smoke evaluation today returned no scores across all five dimensions—execution, grounding, judgment, integrity, and communication—due to API timeouts, resulting in removal from the main leaderboard. The anomaly points to interface-level failure rather than model capability degradation.
On 2026-08-04, the YZ Index Smoke quick-test covered 9 models, with Gemini 2.5 Pro topping the daily ranking at 89.56 points. Smoke is a daily 10-question quick test designed for short-term signal monitoring and does not carry the same weight as the Full weekly leaderboard.
On 2026-08-03, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 ranking first at 95.19 points. Smoke is a daily 10-question quick test for observing short-term signals, not equivalent to the Full weekly ranking conclusions.
GLM-4.6 scored 74.00 on the main leaderboard in today's Smoke evaluation, with 82.30 on code execution and 95.00 on material constraint. Two dimensions are missing due to API failure/timeout and have entered automatic retesting, excluded from this period's ranking.
In the July 28–August 2, 2026 Smoke evaluation, Qwen3 Max posted the largest seven-day gain (+36.8) to close at 96.1, while Gemini 3.1 Pro fell 5.6 points to 94.45, making it the biggest loser. The report also analyzes scoring trajectories, volatility drivers, integrity-rating changes, and implications for users.
On August 2, 2026, the YZ Index Smoke quick test covered 10 models, with Doubao Pro ranking first at 96.7 points. The briefing covers daily scores across code execution and material constraint dimensions, along with key fluctuations and integrity signals to monitor.
MLPerf Endpoints v0.7 marks the foundational release of a buyer-centric AI inference benchmark, publishing initial results from Coreweave, Google, Intel, KRAI, and Nvidia across three benchmarks. The release supports automated submission pipelines, continuous review tooling, and dynamic result visualization, with v1.0 planned for later this year.
SGLang and Miles Add Day-0 Support for Kimi K3SGLang TeamJuly 27, 2026We are excited to announce Day-0 support for Kimi K3 in SGLang and Miles. K3 is the first open-source model in the 3-trillion-para
Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in MilesZiang Li, humans& and Miles TeamJuly 29, 2026 TL;DR: We implemented two Blackwell-native RL recipes in Miles: end
RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUsRadixArk & GoogleJuly 30, 2026RadixArk and Google Cloud are partnering to bring SGLang to TPUs, giving developers ultimate fl
Toward a Cleaner Quantization Stack in SGLangSGLang X Ascend TeamJuly 28, 2026Quantization has moved from an advanced feature to an essential part of high-throughput LLM serving. As the number of chec
In today's Smoke evaluation, GLM-4.6's material constraint score dropped from 75.00 to 47.70 points, while its main score rose from 46.29 to 76.47 points.
GPT-o3 scored 79.28 points on today's Smoke evaluation main leaderboard, down 13.9 points from yesterday's 93.16, with notable declines in both code execution and material constraint dimensions.
On 2026-08-01, the YZ Index Smoke quick test covered 11 models, with Claude Opus 4.7 and Qwen3 Max tying for first place at 93.39 points. Key signals include GLM-4.6's integrity dropping to warn and multiple models posting sharp overall declines.
In today's Smoke evaluation, Qwen3 Max's material constraint score dropped 20 points to 47.70, while code execution soared 37.8 points to 92.50, lifting the main leaderboard by 11.8 points to 72.34.
In today's Smoke evaluation, Grok 4's code execution score dropped from 92.00 to 72.50, while material constraint rose from 60.90 to 84.10, and the main leaderboard score slightly fell from 78.01 to 77.72.
On 2026-07-31, the YZ Index Smoke quick test covered 10 models, with DeepSeek V4 Pro scoring 96.94 to top the daily rankings. The test focuses on code execution and material constraints, serving as a short-term signal indicator.