Around August 20, 2026, an unknown model named Ox Alpha appeared on the OpenRouter platform under the stealth/ox-alpha designation, supporting a 1.05 million token context window and multimodal input across text, image, and video. Early tests show it outperforms GPT-5.6 Sol and Fable 5 on coding and agent tasks, as confirmed by independent media.
Fact Reconstruction
According to OpenRouter page information, Ox Alpha was released around August 20, offering one week of free usage, with a context window of 1.05 million tokens, a maximum output of 130,000 tokens, and support for text, image, and video input. Source 2 shows its cache hit rate is approximately 77.9%, with the knowledge cutoff extending to November 2025. Source 3 confirms it is the first model in the Stealth anonymous series to natively support full modality, currently available for free on OpenRouter and OpenCode Zen, operating under a zero data retention protocol.
Mechanism Breakdown
The model employs an anonymous deployment strategy, with no company announcement or brand promotion, accessible only through the OpenRouter platform. Source 2 notes that its tokenizer characteristics match the Zhipu GLM series across all 11 metrics, consuming 29 tokens in digital probe tests—a perfect match with GLM rather than DeepSeek's 98 tokens. The deployment configuration is highly consistent with GLM-5.3, including tightened parameter interfaces, and Ox Alpha adopted the same settings just two days after GLM-5.3's new configuration went live. Zhipu founder Tang Jie previously stated that a multimodal mode was coming soon, and three months have passed since that statement.
Industry Impact
For developers, the 1M context allows direct input of an entire codebase for refactoring without RAG chunking, and video input can be used to diagnose frontend interaction bugs. Enterprise users can quickly integrate tools such as Cursor and Aider through the standard OpenAI interface, with the zero-retention protocol reducing data risk. For the competitive landscape, the anonymous testing model allows Zhipu to collect high-concurrency data before official release, following a path similar to that of GLM-5 and MiMo-V2-Pro. For upstream and downstream platforms, OpenRouter gains traffic growth, with Source 3 indicating it provides ultra-high concurrency service capabilities.
Comparison and Precedents
Source 3 comparison shows that Ox Alpha's 1M context is close to that of Gemini 3.7 Flash, but it also supports video input and is currently free; DeepSeek V4 Flash has a context of only 128K and uses peak-valley billing. In historical Stealth cases, anonymous models were ultimately revealed to be from the GLM-5 or MiMo series, and both Ox Alpha's Tokenizer match and its GLM-5.3 configuration consistency point to the same origin.
Strategic Assessment
Based on existing facts, Zhipu is most likely to officially release a multimodal GLM version in the near future, with Ox Alpha serving as a warm-up test that has already validated 1M context and video capabilities. Developers can test its long-context coding performance through the OpenRouter API. The above assessment is based on publicly available Tokenizer probe and configuration matching analysis.
For selection recommendations, developers who need to process entire codebases and multimodal input can prioritize trying Ox Alpha's free quota while keeping GPT-5.6 Sol as a baseline for comparison; enterprises should monitor the actual implementation of the zero-retention protocol and evaluate long-term pricing after the official version is released. Source 1 and Source 3 both confirm its current pricing is $0 per million tokens.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接