Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deployments
Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE DeploymentsThe Mooncake Team, Volcano EngineMarch 25, 20261.
Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE DeploymentsThe Mooncake Team, Volcano EngineMarch 25, 20261.
The MLPerf Inference v6.0 release marks a significant expansion in open-source large language model (LLM) coverage, introducing two key additions to the Reasoning LLM task group: the GPT-OSS 120B benchmark based on a high-capacity MoE model, and a new interactive workload for DeepSeek-R1 with low-latency constraints, featuring the first standardized speculative decoding in MLPerf.
GPT-4o suffered a dramatic 10.3-point drop in Material Constraints this week, falling to last place among 11 models, while Baidu's Wenxin 4.0 became the only model to achieve positive growth in core dimensions.
Doubao Pro's stability score crashed from 54.5 to 34.7 in the latest YZ Index evaluation, revealing serious consistency issues that could undermine user trust and enterprise adoption.
ROCm Support for Miles: Large-Scale RL Post-Training on AMD Instinct™ GPUsAMD & Miles TeamMarch 17, 2026Reinforcement learning (RL) has rapidly become a core stage of modern foundation-model development.
Grok 3's stability score crashed from 54.2 to 31.7 points in the latest Winzheng evaluation, exposing a fatal weakness in current AI models that excel at coding but fail at real-world engineering judgment.
GPT-o3's availability score plummeted from 100 to 69 in just one week, exposing fundamental architectural defects rather than isolated issues—a technical accident that reveals systemic imbalances in AI development.
GPT-o3 suffered a catastrophic system failure with stability plummeting from 53 to 28 points and availability dropping from 100 to 69, revealing fundamental architectural flaws rather than typical performance variations.
GPT-o3's long context processing capability collapsed in recent testing, with scores dropping from 62.3 to 28.8 points due to aggressive API rate limiting, exposing serious infrastructure issues at OpenAI.
GPT-4o experiences a catastrophic performance collapse with its usability score plummeting from 100 to 65, caused by overly conservative "strict tool calling" that makes the model refuse to perform basic tasks.
Doubao Pro's stability score plummeted from 54.5 to 34.7 (a 36.3% drop) this week, despite significant improvements in programming and knowledge work dimensions, revealing a concerning pattern of "progress and regression coexisting" that warrants in-depth analysis.
GPT-4o's catastrophic failure in long-context tests, with 5 questions returning rate limit errors, reveals OpenAI's severe infrastructure problems rather than model capability issues.
Gemini 2.5 Pro's stability score plummeted 22.8 points in one week, exposing a critical lack of engineering judgment despite gains in programming capabilities.
Wenxin 4.0's stability score crashed from 52.1 to 30 points while programming ability soared by 41.4 points, exposing Baidu's critical engineering shortcomings and raising serious concerns about China's AI industrialization approach.
Qwen Max exhibits extreme duality in this week's evaluation, with significant improvements in programming and long-context tasks, but a catastrophic decline in stability metrics. This "fire and ice" performance warrants in-depth analysis.
This week's evaluation data reveals Gemini 2.5 Pro's stability score plummeted from 54.0 to 31.2, a 42.2% drop, exposing serious issues in maintaining consistent output quality while other metrics improved.
DeepSeek R1's stability score crashed from 53.7 to 31.6 points this week, with the model failing basic judgment questions like whether water can boil at 101°C under standard pressure, raising serious concerns about its reliability.
While everyone celebrates Claude's 38.3-point programming improvement, a more dangerous signal has been masked: stability plummeted from 54.2 to 31.2 points, revealing a systemic algorithmic collapse rather than normal performance fluctuation.
Wenxin Yiyan 4.0 showed remarkable anomalies in this week's evaluation, with programming capability surging 41.4 points but stability plummeting from 52.1 to 30.0 points, revealing potential deep-seated issues in the model upgrade process.
DeepSeek V3 shows contradictory performance this week with programming capabilities soaring 42.6 points while stability metrics collapse from 53.4 to 32.0 points, revealing critical trade-offs in AI model optimization.