Naive-N0.5-Flash: the 309B open-weight MoE with no full-attention layers at all aims for 2,000 tokens/s per user
🔎 309 billion parameters, zero full-attention, and an awkward revelation
On September 27, 2026, NaiveAI — a Beijing-based startup founded by Jifeng Dai, a researcher from Tsinghua, valued at $1.4 billion — published on Hugging Face the weights of Naive-N0.5-Flash, a 309-billion-parameter MoE (15.5 billion active), under an MIT license. In an era of weekly open-weight releases, nothing extraordinary on the surface. Except for one detail: of its 48 layers, none is a full-attention layer.
Instead: 39 sliding-window layers and 9 DeepSeek Sparse Attention layers, a native one-million-token context, and a homegrown inference stack called NaiveRT that claims up to 2,000 tokens/s per user. According to AI Weekly, this is the first frontier-scale open-weight release with no full-attention layers at all.
And then there's the most troubling sentence in the write-up: the company claims that its R&D pipeline was "substantially executed by AI systems," with humans only setting objectives and evaluation standards. This launch therefore tests two bets at once — one architectural, the other organizational. Both deserve a closer look.
The Essentials
- Radical architecture: MoE with 309 billion parameters (15.5 billion active), 48 layers — 39 SWA + 9 DSA — with no full-attention at all, native 1M-token context.
- Documented training recipe: 3.25 T tokens in three stages, on Xiaomi's open-weight MiMo-V2.5 base; weights and inference code under MIT license.
- NaiveRT, the in-house inference stack: 50 tokens/s per user in Standard mode, up to 2,000 in Ultrafast (fused mega-kernel, PDL, speculative decoding).
- API announced at $0.10 / $0.40 / $0.01 per million tokens (input / output / cache reads) — but not yet live at the end of September 2026.
- The real story: R&D that was "substantially executed by AI systems," and a NaiveRT built in six days across 151 documented optimization attempts.
- Our take: the speed figures are real but narrow (decoding peaks, not end-to-end benchmarks). The deeper issue is AI-assisted R&D velocity.
Recommended Tools
| Tool | Main use | Price | Ideal for |
|---|---|---|---|
| Naive-N0.5-Flash weights (Hugging Face) | Download the weights (FP8, ~315 GB) | Free — MIT | Labs, self-host teams |
| Inference code (GitHub) | Run and dissect the inference | Free — MIT | Inference engineers |
| NaiveAI research page | Detailed NaiveRT specs and benchmarks | Free | Understanding the optimizations |
| NaiveAI API (upcoming) | Managed access to the model | $0.10 / $0.40 / $0.01 per M tokens (announced, September 2026, check naive.ai) | Devs without a GPU cluster |
| Hostinger | Host the application that will consume the API | From a few €/month (October 2026, check hostinger.com) | Building a product around the model |
Zero full-attention: what it really changes (and what it doesn't)
By removing all full-attention layers, NaiveAI mainly reduces the computational cost of attention — not memory growth, which continues to climb with context length. That's the whole subtlety, and it matters.
Concretely, the architecture boils down to a few numbers, all detailed on the Hugging Face model card: 48 layers, including 39 using sliding-window attention with a window of 128 tokens, and 9 using lightweight DeepSeek Sparse Attention (GQA4), at a ratio of roughly 5:1. The DSA does not scan the entire history: a lightweight indexer with 16 query heads selects the 2,048 most relevant tokens for the backbone attention, with a sink bias applied to both attention types.
The bet is clear: handle 1M tokens of native context without ever paying the quadratic price of full attention. The training recipe is also spelled out: 3.25 T tokens in three stages — 50B tokens of Indexer Warmup (a KL loss aligns the indexer with the attention of converted layers), 3T tokens of Sparse Attention Training as continued pretraining focused on code and AI research, then 200B tokens of Learning Rate Decay. All built on Xiaomi's MiMo-V2.5 base, itself open-weight.
But let's read the footnote highlighted by the critical analysis from canberk.me: "no full attention" does not mean flat memory usage. The indexer keeps scanning the history, and the complete KV cache is retained. Sparse attention reduces attention compute and memory accesses — not the amount of state that must be stored. If you're sizing an infrastructure, this detail is what counts.
This movement isn't happening in isolation. DeepSeek paved the way with its sparse attention — which NaiveAI explicitly credits — and our coverage of DeepSeek V4.1 Flash under MIT license already showed a KV cache reduced by a quarter. More recently, Subquadratic came out of stealth with SubQ and 12 million tokens of context, pushing the logic to its conclusion. My take: NaiveAI's hybrid SWA + DSA is the most pragmatic compromise available right now — sliding-window for nearby context, DSA for the long term, and no layer forced to traverse an entire million tokens anymore.
NaiveRT: 2,000 tokens/s, yes — but read the footnote
NaiveRT's numbers are real and documented, but they are narrow decoding measurements, not guaranteed production throughput. The distinction is essential before any capacity planning.
The in-house stack claims 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast, via three levers: mega-kernel fusion, Programmatic Dependent Launch (PDL), and speculative decoding. The peak figure highlighted by naive.ai: 2,122 tokens/s in single-stream decoding on 8 GPUs — best one-second window, thinking off, 41 HTML/SVG generation requests, prefill excluded, temperature 0.4. In other words, a timed sprint under ideal conditions, not a marathon.
Other metrics are more telling for specialists: a full speculative round in 3.4 ms versus 12.3 ms for SGLang on the same system, and a lightweight indexer that cuts selection wall time by 44% compared to the original DSA implementation. Pandaily confirms the release's overall coherence.
The use case that justifies all this is RL. With 1M-token contexts, a rollout's wall-clock time is dominated by the tokens generated — 50 to 100 tokens/s in conventional decoding. Hence NaiveAI's business argument: fast single-stream decoding directly transforms training throughput. Supporting evidence: with the 1M-context configuration, 1T training tokens processed in ~4 days on 512 GPUs. For teams doing RL at scale — a topic we cover in detail in our guide to the best LLMs for AI agents — that's an argument that holds up.
RuntimeWire notes this explicitly, and AI Weekly echoes it: the 2,000 tokens/s are narrow decoding-speed figures, not end-to-end benchmarks. canberk.me adds that it's neither sustained throughput nor generic API speed.
My take: don't judge the model on these numbers, judge the stack. Publishing inference code under MIT with detailed measurement conditions is more honest than the industry average — and it's precisely what enables this kind of critical reading.
"Substantially executed by AI systems": the real story
Beyond the architecture, NaiveAI's most important claim is organizational: its R&D was allegedly carried out mostly by AI systems. That may be the real product of this release.
The Beijing-based startup, founded by Jifeng Dai (Tsinghua) and valued at $1.4 billion, claims that the model's R&D was substantially executed by AI systems, with humans setting objectives and evaluation standards. Even more concrete: according to Cellcog, NaiveRT was allegedly "built in six days by human researchers working with AI models, across 151 documented optimization trials." Six days. 151 logged trials.
What sets this claim apart from the usual marketing is the documented data: tracked optimization trials, the announced development timeline, published benchmark conditions. It's not an independent audit — far from it — but it's infinitely more verifiable than the usual press releases about "AI doing research." RuntimeWire confirms that the company fully embraces this narrative, starting from Xiaomi's open MiMo-V2.5 base and crediting DeepSeek's sparse attention work.
AI Weekly sums up the double bet well: sustaining 1M context without full attention, and proving that AI tooling materially accelerates a frontier lab's R&D. With 15.5 billion active parameters, the per-token compute is moreover close to that of a mid-sized dense model — efficiency is part of the demonstration.
If the second hypothesis holds at scale, the labs' competitive advantage will no longer be cluster size, but the quality of their research agent pipelines. It's a race China seems determined to lead: Moonshot AI just raised $2 billion with Kimi K2.6 at the top of open-weight. The usual caution still applies: "substantially" has no precise definition, and no third party has audited anything yet.
Pricing, hardware, access: what it takes to get your hands on it
The model is free in theory (MIT), but running it requires a multi-GPU node; the managed API is not yet available. Here's a blunt summary.
On the hardware side, expect roughly 315 GB of weights in FP8, an FP8-compatible NVIDIA GPU is mandatory, and note that the 2,122 tokens/s peak was measured on 8 GPUs. We're a long way from the "local LLM on a gaming PC" — for that use case, our local LLM installation guide and our selection of the best LLMs to run locally remain the right starting points.
On the API side, here are the announced prices:
| Item | Announced price | Status (end of September 2026) |
|---|---|---|
| Input tokens | $0.10 / M tokens | Announced, not yet live |
| Output tokens | $0.40 / M tokens | Announced, not yet live |
| Cache reads | $0.01 / M tokens | Announced, not yet live |
At the time of checking Cellcog on September 27, 2026, the API was announced but not live, with no OpenRouter listing. Practical consequence: don't budget anything firm before it actually goes live, and check the prices on naive.ai when making your decision. If you're already preparing the application layer that will call this API, a standard hosting solution like Hostinger is more than enough — it's the model that's heavy, not your frontend.
Facing open-weight: Naive plays a different tune
NaiveAI doesn't (yet) compete on quality benchmarks — it competes on architecture, inference speed, and the narrative of automated R&D. The positioning is different from its Chinese neighbors.
| Model | License | Key takeaway |
|---|---|---|
| Naive-N0.5-Flash (NaiveAI) | MIT (weights + inference) | 309B (15.5B active), 1M native, zero full-attention, targeted at code and AI R&D |
| DeepSeek V4.1 Flash | MIT | 552B, KV cache reduced by a quarter |
| MiniMax M3 | Open weights | 1M context, positioned against GPT-5.5 |
| GLM-5.3 (Z.AI) | Open weight, anti-hyperscaler license | Openness with restrictive conditions |
| Kimi K2.6 (Moonshot AI) | Open weight | Leader of open-weight agentic; Moonshot raised $2B |
Important journalism point: no independent public quality measurements accompany the release. The model is positioned for code and AI R&D, which theoretically puts it in competition with Claude Opus 4.7, GPT-5.5, or Gemini 3 Pro — our comparison of the best LLMs for coding will be updated as soon as community evaluations come in. Until then, any claim of superiority would be puffery.
In any case, the open-weight landscape is diversifying at a record pace: Z.AI with GLM-5.3 going open-weight under an anti-hyperscaler license, MiniMax with M3 and its one million tokens of context, and even specialized niches like Cohere's open-weight translation MoE that beats DeepL and Google Translate on WMT26. NaiveAI adds to this wave a piece nobody had: an architecture without full attention at frontier scale.
My take: the value of this release is primarily strategic. It forces everyone to admit that the architectural consensus — full attention, growing KV cache, MoE — is not a terminus. That's good for the entire ecosystem.
Verdict: who should pay attention now
Inference engineers, RL teams, and agent builders: yes, right away. End users: wait for the API and benchmarks, there's no rush.
Who it's relevant for today:
- Long-context RL teams: the fast single-stream decoding argument is the most concrete part of the entire release, backed by training throughput figures.
- Systems engineers: the inference code on GitHub is a goldmine — mega-kernels, PDL, speculative decoding benchmarked against SGLang, all under MIT.
- Efficiency researchers: 48 layers without full attention, a 16-head indexer, a fully documented multi-stage training recipe.
Who it's not (yet) the right time for: developers who want a reliable everyday coding model will stick with the trusted picks in our comparison of the best LLMs for coding, and free general-purpose use is well covered by our selection of the best free LLMs. For an overview of the market, our monthly comparison of the best LLMs sorts it all out.
To watch for: the API going live, a possible OpenRouter listing, and above all the first community feedback on the model's actual coding quality.
❌ Common Mistakes
Mistake 1: believing that "zero full-attention" means constant memory
This is the most widespread misunderstanding. The indexer keeps scanning the history and the full KV cache is retained: memory still grows with context. Sparse attention reduces attention compute and memory accesses, not the state that has to be stored. Solution: size your VRAM for the KV cache at 1M tokens, not for the 315 GB of weights alone.
Mistake 2: taking 2,122 tokens/s as a guaranteed throughput
This figure is a single-stream decoding peak: best one-second window, prefill excluded, 8 GPUs, thinking off, temperature 0.4. It is neither a sustained end-to-end throughput nor a generic API speed. Solution: wait for end-to-end measurements and real API throughputs before sizing anything.
Mistake 3: wanting to run it on your dev machine
315 GB of weights in FP8, an FP8-compatible NVIDIA GPU required, benchmarks on 8 GPUs: this is not a consumer-grade model. Solution: go through the API once it's live, or pick a model that can genuinely be run locally via our dedicated local guides.
Mistake 4: building a product on an API that doesn't exist yet
The API was announced but not live as of late September 2026, with no OpenRouter listing. Building on it today is building on sand. Solution: abstract your provider and plan for an immediate fallback — DeepSeek V4 Pro or Kimi K2.6 do the job very well in the meantime.
❓ Frequently Asked Questions
Is Naive-N0.5-Flash really free of any full-attention layers?
Yes. The Hugging Face model card details 48 layers: 39 with sliding-window attention (128-token window) and 9 with DeepSeek Sparse Attention, with no full-attention at all. According to AI Weekly, it's the first open-weight release at frontier scale with this characteristic. Note however: KV cache memory keeps growing with context.
Can you run it locally?
Technically yes — MIT license, with open weights and inference code. In practice, expect around 315 GB of weights in FP8, FP8-compatible NVIDIA GPUs, and a multi-GPU node (NaiveAI's benchmarks run on 8 GPUs). This isn't a "gamer PC" model; for accessible local setups, follow our local LLM installation guide.
How much will the API cost?
NaiveAI announces $0.10 per million input tokens, $0.40 for output, and $0.01 for cache reads (September 2026, check naive.ai). But the API wasn't live yet as of September 27, 2026, and there's no OpenRouter listing. Wait for the actual launch before budgeting a project.
How good is it for code compared to Claude Opus 4.7 or GPT-5.5?
Impossible to say today: no independent public benchmarks accompany the release. The model is positioned for code and AI R&D, built on a MiMo-V2.5 base trained on 3.25 T tokens, but the quality remains to be demonstrated. Our comparison of the best LLMs for coding will be updated as soon as the first community feedback comes in.
Is the claim of "AI-executed" R&D credible?
Partially verifiable. NaiveAI documents 151 optimization trials and a NaiveRT built in six days, with published benchmark conditions — more traceability than the industry average. But "substantially executed by AI systems" remains a corporate claim, without independent audit. Treat it as a serious hypothesis, not an established fact.
✅ Conclusion
Naive-N0.5-Flash may not change your day-to-day life as a developer this week, but it could change what labs believe is possible — architecturally, with this full-attention release, and organizationally, with R&D driven by AI systems. Keep an eye out for the API launch and, in the meantime, check out our monthly comparison of the best LLMs to pick a model that's actually available today.