On September 27, 2026, Beijing-based lab NaiveAI released the 309B MoE open-source model Naive-N0.5-Flash, with 15.5B activated parameters, support for 1M context, full open-source release under the MIT license, and inference speeds of up to 2000 tok/s.
Factual Reconstruction
The model was launched by Beijing-based lab NaiveAI on September 27, 2026. It has 309B total parameters, 15.5B activated parameters, and native support for 1M context. Architecturally, it has no global attention layers; it uses a hybrid design combining sliding-window attention and DeepSeek sparse attention, is adapted from Xiaomi MiMo-V2.5, and was continually trained on 3.25T tokens. The weights and inference code are open-sourced under the MIT license, positioned for coding and AI R&D tasks.
Mechanism Breakdown
The training process uses AI participation in optimization and R&D. AI models are responsible for exploring hybrid attention architectures while optimizing training, inference, and deployment systems. Human researchers provide directional guidance and key decisions. The inference system NaiveRT combines mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding, delivering 50 tokens/s per user in Standard mode and up to 2000 tokens/s in Ultrafast mode.
API pricing is $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached read tokens. The model also offers open weights to developers.
Industry Impact
For developers, the MIT license allows direct downloading of weights and code modification, while the 1M context and 2000 tok/s peak speed lower the barrier to long-text processing and rapid iteration. For enterprise users, the API pricing offers a quantifiable option for invocation costs.
For the competitive landscape, the model's hybrid attention design distinguishes it from existing full-attention models, and developers can use it to test the real-world performance of local or sparse computation in coding tasks. For upstream and downstream toolchains, NaiveRT's optimization path may influence kernel fusion and scheduling strategies in future inference frameworks.
Comparisons and Precedents
The model was continually trained on 3.25T tokens, and its architecture was adapted from Xiaomi MiMo-V2.5. The comparison models mentioned in evaluations include GLM-5.3, Kimi-K3, Qwen-3.8-Max, and others, but specific scores should be referenced from each model's official blog.
Strategic Assessment
Based on current facts, the most likely next development is that the developer community will conduct fine-tuning and tool integration around the MIT-licensed weights; signals to watch include GitHub commit counts, real-world inference latency reports, and changes in API call volume.
When developers are selecting models, they can first deploy the weights locally to test coding tasks under 1M context, then compare API costs with the cost-effectiveness of self-hosted inference. Enterprise users can validate the stability of Ultrafast mode under production loads through small-scale API calls before deciding whether to expand usage.
The recursive self-improvement path claimed by the model is still at an early stage. Its advantages lie in fully open-source weights and code, as well as its training positioning for AI R&D tasks.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接