Two Open-Source Surprise Strikes: Alibaba's 2.4-Trillion-Parameter Model and Zhipu's Mysterious Powerhouse Unveiled in the Same Week

In the final week of August 2026, Alibaba's Qwen team open-sourced the 2.4-trillion-parameter Qwen3.8-Max, while Zhipu AI's Z.ai revealed that the mysterious "Ox Alpha" model on OpenRouter was actually GLM-5.3-Flash. The twin releases highlight contrasting open-source strategies, a 26-fold inference cost gap, and a shared bet on coding and agent workloads.

Qwen GLM Open Source AI
645

Zhipu Open-Sources GLM-5.3-Flash: Anonymous Model Ox Alpha Tops Rankings in Six Days, Matches Frontier Scores at One-Fortieth of Opus 4.8's Price

Zhipu AI officially open-sourced GLM-5.3-Flash (320B-A18B) on August 26, 2026, confirming it was the anonymous model "Ox Alpha" that topped call-volume rankings on OpenRouter and OpenCode within six days. Matching Claude Opus 4.8's frontier score at roughly one-fortieth the price, the model completed its entire anonymous testing phase on domestic Chinese chips.

智谱AI GLM-5.3-Flash Ox Alpha
696

One Video Teaches Robots 10-Minute Tasks: Skild AI Releases S1 Foundation Model, Breaking the Robot Training Paradigm

Skild AI has released the S1 robot foundation model, which learns multi-step physical tasks of up to 10 minutes from a single human demonstration video—without fine-tuning or post-training—achieving a 66% success rate on unseen tasks versus 9% for language-prompted VLA models. The approach treats video demonstrations as programs, leveraging trillion-scale simulation pretraining for in-the-wild generalization.

Robotics 基础模型 Skild AI
730

AWS and NVIDIA Add 2 Million GPUs: Demand Exceeds Every Forecast, Partnership Expands from Chips to Full Technology Stack

On August 26, 2026, AWS and NVIDIA jointly announced the deployment of an additional 2 million NVIDIA GPUs across AWS's global infrastructure from 2027 to 2028, spanning the Blackwell Ultra, Rubin, and Rubin Ultra architectures. Driven by customer demand that has exceeded all previous forecasts, the partnership has expanded from GPU procurement to full-stack co-design across CPUs, networking, memory, open models, and robotics.

AWS NVIDIA GPU
841

OpenAI's Self-Developed Chip Jalapeño Passes First Test: Per-Watt Performance Up to 1.9x NVIDIA's Flagships, 700W vs 1400W

On August 26, 2026, OpenAI released the first performance data for its self-developed inference chip, Jalapeño, showing 1.5-1.9x higher throughput per kilowatt and 1.7-3.6x lower end-to-end latency than NVIDIA GB200 and GB300 rack systems on the InferenceX benchmark, marking the company's formal entry into hardware-level competition.

OpenAI Jalapeño 自研芯片
296

Nvidia Teams Up with Six Wall Street Giants to Raise Over $500 Billion: Computing Power Is Turning Into a Bond

Nvidia has joined forces with six top global asset management institutions to establish an independent computing power financing platform targeting over $500 billion in third-party capital. The move aims to reposition AI compute as an investable asset class, but critics warn of structural risks including GPU depreciation mismatches and the use of pension funds for AI infrastructure financing.

NVIDIA AI Infrastructure 华尔街
553

Google Gemini 3.5 Transcribe Opens API: 2.6% Word Error Rate Ranks Fifth, Real-Time Streaming Has Three Hard Limits

Google officially released Gemini 3.5 Transcribe via the Gemini API on August 26, 2026, achieving a 2.6% word error rate (ranked fifth) in non-streaming benchmarks and 4.0% in real-time streaming. The release splits into batch and streaming endpoints with three hard constraints, while Google's broader strategy extends from API access to embedding voice input directly into Chrome.

Google Gemini 语音识别
742

OpenAI 'Jalapeño' Chip Benchmark Debut: 700W Processor Outperforms Nvidia's 1400W Flagship, Inference Landscape Begins to Shift

OpenAI unveiled the first public benchmark results for its self-developed AI inference chip, Jalapeño, at Hot Chips 2026, delivering 1.5–1.9x the per-watt inference throughput of Nvidia GB200/GB300 systems with end-to-end latency cut to 28–59% of comparison systems. The 700W chip comprehensively outclasses Nvidia's 1400W flagship in efficiency.

OpenAI Jalapeño芯片 NVIDIA
719