NaiveAI has released Naive-N0.5-Flash on Hugging Face, Pandaily reported on 28 September. The open-weight model is aimed at coding and AI research and development, with an MIT licence covering its weights and inference code.
The model has about 309 billion parameters in total, but activates 15.5 billion for a given task. That is the principle of its mixture-of-experts design: instead of using the whole network for every piece of text, it draws on selected parts. Its stated native context window is one million tokens, allowing it to take in an unusually long stretch of text at once.
To handle that length, NaiveAI combines sliding-window attention, which concentrates on nearby text, with lightweight DeepSeek Sparse Attention, which selects information from farther away. It describes a roughly 5:1 arrangement of the two methods, alongside grouped-query attention, with no full-attention layers.
Reviewing a long collection of code could involve fewer separate excerpts if the material fitted within that context window. Whether the model would find the relevant parts and use them accurately is a different test from how much text it can accept.
Naive-N0.5-Flash builds on the open-weight MiMo-V2.5 base through further mid-training and post-training, rather than full pretraining from scratch. For NaiveRT, NaiveAI puts Ultrafast mode at up to about 2,000 tokens a second and Standard mode at about 50 tokens a second for each user.