The 2-bit revolution: Ternary-Bonsai-2 and 27B GGUFs run multimodal models on a consumer GPU
🔎 A week that shifted the landscape
While most observers were scrutinizing the scores of the cloud giants, HuggingFace lived a different story this week: one of extreme compression. Ternary-Bonsai-2-27B, a 27-billion-parameter multimodal model distributed in 2 bits, surpassed 4 million downloads (October 2026, agents-radar). In the same window, ISTA-DASLab published four GSQ-RCO quantizations in seven days, and an uncensored GGUF crossed the 1.5 million download mark.
Three numbers, one signal: local AI on consumer hardware is no longer the exception reserved for RTX 4090 owners. It's the emerging norm, week after week, in the download rankings.
What makes this wave different from previous ones isn't the marketing. It's the engineering: for the first time, a native 2-bit model delivers 98.2% of the quality of its FP16 base, where a classic 2-bit quantization caps out at 84.1%. We break down why, what it concretely changes, and what you actually lose.
The essentials
- Ternary-Bonsai-2-27B: 27.36B multimodal parameters in 5.95 GB (PTQ1_0, 1.75 bit/weight), versus ~54 GB in FP16 — a 9.3× reduction.
- 98.2% of FP16 quality (84.78 vs 86.32 on the thinking average), nearly on par with a 4-bit UD and far ahead of a classic 2-bit quant (84.1%).
- Two packings: PTQ1_0 (dense trits, king of prompt processing) and PQ2_0 (2-bit slots, faster at decoding on Ada/L4 and in tight memory situations).
- Mandatory fork: the ternary kernels live in the PrismML-Eng fork of llama.cpp; the stock binary produces garbage.
- Mac served: official MLX variant, 4-bit KV cache, DSpark drafter for speculative decoding. Apache 2.0 license, CUDA/Metal/CPU backends.
Recommended Tools
| Tool | Main use | Price (October 2026) | Ideal for |
|---|---|---|---|
| Ternary-Bonsai-2-27B GGUF (PTQ1_0) | Multimodal 27B, 262K context | Free (Apache 2.0) | 8 GB+ GPUs, max decoding on H100/A100/Blackwell |
| Ternary-Bonsai-2-27B GGUF (PQ2_0) | 2-bit slots variant (7.21 GB) | Free (Apache 2.0) | RTX Ada / L4, tight VRAM |
| MLX 2-bit version | Apple Silicon (Python/Swift) + CUDA | Free | 32 GB M-series Macs |
| Fork llama.cpp PrismML-Eng | Ternary CUDA/Metal/CPU kernels | Free (open source) | Serving the model locally |
| Hostinger | Hosting the interface or API that exposes your model | From ~€3/month (check hostinger.com) | Putting your instance into production |
Why 2-bit Finally Works
2-bit works now because the compression is designed into the model, not applied on top of it like a band-aid. Three ingredients, documented on the model's official card:
Ternary weights with FP16 scales. Each group of 128 weights (g128) contains only three values — −1, 0, +1 — but keeps its scale in FP16. The trit carries the direction, the scale carries the magnitude. It's this duo that makes it possible to get down to 1.75 bits per weight without murdering the information.
Hadamard rotations in blocks of 1024. Outliers, the historical enemies of low-precision quantization, are dispersed before compression instead of being brutally clipped. That's the difference between 98.2% and 84.1% of FP16 quality.
Hybrid attention, ~75% linear / ~25% full. Only 16 of the 64 layers carry full attention. Fewer layers to memorize carefully, hence a well-controlled KV cache — and a 262K-token context that becomes realistic on a personal machine.
We charted this trajectory in our feature on 1-bit LLMs, when models fit on a smartphone. What changes with Bonsai 2 is the proof at scale: a complete 27B multimodal model, published benchmarks to back it up, not just a simple lab demo.
One implementation detail is worth highlighting: the weights are never re-expanded to FP16. The fork's kernels compute directly on the packed trits. Now, decode-time inference is limited by memory bandwidth — dividing by nine the data to be moved mechanically divides the bottleneck. That's why the gains are most visible precisely where memory is scarce.
What's Really Inside Ternary-Bonsai-2-27B
A full multimodal model — 27B backbone, vision, 262K context — in the footprint of an AAA video game.
The base is Qwen3.8-27B, architecture unchanged. This point matters: this is not a model trained from scratch in ternary, but a controlled conversion of a foundation whose quality is already established. Native quantization applies to the weights, not the architecture.
The detailed bill (source: HuggingFace repository, October 2026):
| Component | Parameters | Role |
|---|---|---|
| Backbone (64 blocks) | 24.35 B | Language and reasoning, hybrid attention |
| Embeddings + LM head | 2.54 B | Vocabulary |
| Vision tower | 0.46 B | Image understanding |
| Total | 27.36 B | Qwen3.8-27B base |
Two GGUF packings are offered, and the choice is not trivial:
| Packing | Bits/weight | Size | Strength |
|---|---|---|---|
| PTQ1_0 | 1.75 | 5.95 GB | Dense trits; fastest prompt processing across all backends |
| PQ2_0 | 2.13 | 7.21 GB | 2-bit slots; fastest decoding on Ada/L4 and in tight memory |
To put things in perspective: the announced sweet spot is 1.72 bpw, i.e. 5.8 GB — roughly a 9.3× reduction compared to the ~54 GB of FP16. A model that used to require a datacenter server now fits, moderate context included, on an 8 GB card. This is not incremental progress, it's a step change.
Quality: 98.2% of FP16 at 1.72 bit
On the average of the thinking benchmarks, Bonsai 2 gives up 1.5 points compared to FP16 — and opens up a 12-point gap with classic 2-bit quantization.
The measurements published on the repository (October 2026):
| Quantization | Bits/weight | Thinking score (average) | % of FP16 |
|---|---|---|---|
| Qwen3.8-27B FP16 (reference) | 16 | 86.32 | 100% |
| UD-Q4_K_XL | 4 | 85.18 | 98.7% |
| Bonsai 2 27B (native ternary) | 1.72 | 84.78 | 98.2% |
| IQ2_XXS (classic 2-bit quant) | 2.16 | 72.59 | 84.1% |
The key point of this whole news item fits in one line: a native ternary 2-bit is virtually equivalent to a 4-bit UD, and far ahead of a post-hoc 2-bit quantization. Concretely, the gap between 84.78 and 72.59 separates a model usable on a daily basis from a model that falls apart as soon as the task goes beyond the trivial.
A note on methodology: these figures come from the model's publisher. They are detailed and reproducible, but keep your usual reflex — test on your own use cases. The gap with IQ2_XXS is, however, too wide to be a benchmark artifact.
My take: the lesson is not "2-bit is good". It is that not all quantizations are equal, and that comparisons mixing native quant and imposed quant are misleading. Demand the bits/weight AND the method.
PTQ1_0 or PQ2_0: the right packing for your GPU
H100, A100, or Blackwell → PTQ1_0. RTX Ada or L4 → PQ2_0. And for swallowing long prompts, PTQ1_0 wins everywhere.
The repo publishes measurements per architecture (October 2026):
| Your hardware | Recommended packing | Why |
|---|---|---|
| H100 / A100 / Blackwell | PTQ1_0 | Fastest decode |
| RTX Ada / L4 | PQ2_0 | Fastest decode |
| Tightest memory | PQ2_0 | Best decode throughput, per the repo's measurements |
| Long prompts | PTQ1_0 | Fastest prompt processing everywhere |
The special case of prompt processing
A detail that matters for RAG and long-document analysis: prompt processing favors PTQ1_0 on all tested backends. If your workload consists of ingesting large contexts, stick with PTQ1_0 even on an Ada card, where PQ2_0 nevertheless wins on decode.
The backends cover CUDA, Metal, and CPU — CPU is supported, with honest but unremarkable throughput. On the serving side, a typical llama.cpp launch looks like this (with the PrismML-Eng fork, not the stock binary):
# PrismML-Eng fork required — stock llama.cpp rejects these packings
llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 65536
Adjust the context size according to your VRAM: the 5.95 GB of weights fit in 8 GB, but the KV cache adds on top — see the errors section below.
On Mac: the MLX variant, 4-bit KV cache, and speculative decoding
Yes, it runs on Apple Silicon, and the official MLX variant is the most polished delivery in the lineup.
The MLX version targets Python and Swift, with a CUDA backend as an extra. Three details make the difference in practice:
- Nearly lossless 4-bit KV cache, applied to the full-attention layers (16 out of 64): about 4.3 GB at the full 262K context. It's this figure that makes very long context credible on a personal machine.
- DSpark, the speculative decoding drafter (1.95 GB in Q4_1): a small model proposes tokens, the 27B validates them. On an already fast ternary decoder, the compounded gain is substantial.
- Optional HQQ 4-bit vision tower (~0.63 GB): you only pay for multimodality if you actually use it.
Let's do the budget on a 32 GB Mac: weights (~7 GB), KV cache at the full 262K (~4.3 GB), drafter (1.95 GB), vision (0.63 GB) — everything fits with room to spare for the system and the app. The full multimodal 262K is no longer a theoretical exercise.
We had already seen Qwen3-Coder-Next running on a 64 GB Mac and beating DeepSeek at coding. Bonsai 2 pushes the logic one step lower in memory, and adds vision to the mix.
The week 2-bit took over HuggingFace
This week, extreme compression is no longer a niche — it's the dominant download trend.
The figures compiled by agents-radar (October 2026) give a sense of the phenomenon:
- Ternary-Bonsai-2-27B: 4 M downloads. For a format that still requires a specific fork of llama.cpp, that's a massive vote of confidence from the community.
- ISTA-DASLab: four GSQ-RCO quantizations in one week. An industrial publishing pace, not that of a lone hobbyist.
- An uncensored GGUF: 1.6 M downloads. The "local model without guardrails" segment remains one of HuggingFace's most powerful engines.
- A nesoai mirror already online. When a model is picked up by third parties within days of its release, distribution feeds on itself.
What these numbers tell us is a shift in posture. The community no longer downloads "the biggest model that fits," it looks for "the best model that fits." That's an engineer's reasoning, and it now shapes the rankings.
The movement is consistent at both ends of the chain, moreover. Upstream, DeepSeek is optimizing inter-GPU communication at datacenter scale with DeepEP, its open source library for MoE models. Downstream, extreme compression brings models closer to consumer GPUs. Same philosophy: less waste, everywhere.
For those with no GPU at all, the entry point remains the cloud — our guide on using free models without sacrificing quality covers OpenRouter and Groq. Open heavyweights like DeepSeek V4 Pro (Max) (88 on our index) or Kimi K2.6 (85) remain datacenter territory. But for local use, the 27B class with native quantization is becoming hard to beat — and the frenzy isn't limited to quantizations, with frugal models like Naive N05 Flash confirming that efficiency has become the number one criterion.
What you actually lose — and what you don't
You lose 1.5 points of thinking score, a bit of finesse in vision, and compatibility with stock tooling. You don't lose usability.
The measurable. 86.32 → 84.78 on the thinking average: −1.5 points versus FP16, −0.4 versus the 4-bit UD. In everyday use — writing, code, document analysis — that's below the perception threshold. On reasoning tasks at the model's limits, it can show, and you have to own that.
Vision. The vision tower weighs in at 0.46B parameters: a compact module, not a multimodal giant. For OCR, captioning, UI analysis, the essentials are there. For fine-grained visual understanding — dense diagrams, complex scenes — calibrate your expectations accordingly.
The hybrid architecture. 75% of the attention is linear: that's what makes 262K affordable, but it's a deliberate trade-off on exact recall of information buried far back in the context. For well-chunked RAG, no problem. For retrieving a specific sentence buried at 200,000 tokens, test before you promise.
The ecosystem. No native Ollama or LM Studio support for these packings, no stock llama.cpp: until the ternary kernels are merged upstream, you depend on the PrismML-Eng fork. Historically, this kind of patch does end up merged. But "historically" isn't a date, and in the meantime, your stack needs to know it.
On the other side, the gain is clear: 9.3× compression, a realistic 262K, multimodality included, and an Apache 2.0 license that permits everything — including commercial use. The cost/capability ratio clearly tips in the user's favor.
❌ Common Mistakes
Mistake 1: Loading PTQ1_0/PQ2_0 with stock llama.cpp
Official llama.cpp rejects these formats — and the trap case is worse: it can load a PQ2_0 by treating it as standard Q2_0, producing garbage with no clear warning. Solution: the PrismML-Eng fork, or the MLX variant on Mac. And always check the model's first response on a control prompt before drawing any conclusions.
Mistake 2: Comparing native 2-bit to standard 2-bit quantization
84.78 vs 72.59: the gap says it all. When you're told that "2-bit destroys quality," this usually refers to post-hoc quant without Hadamard rotations or per-group scales. The right comparison is family against family — otherwise the debate is biased before it even starts.
Mistake 3: Forgetting the KV cache in the VRAM budget
5.95 GB of weights doesn't mean "it fits in 6 GB." At full 262K context, the full-attention cache (16 layers out of 64, 4-bit KV) adds about 4.3 GB, on top of which come the drafter and vision tower if you load them. On an 8 GB card: short to moderate context, or partial offload — and there, throughput collapses. Compute the full budget before setting the context size.
❓ Frequently Asked Questions
What graphics card do you need for Ternary-Bonsai-2-27B?
An 8 GB card is enough in PTQ1_0 (5.95 GB of weights) with moderate context. For the full 262K, add the KV cache (~4.3 GB) or move up to a 32 GB Mac. CUDA, Metal, and CPU backends are supported, Apache 2.0 license.
Does 2-bit ternary really degrade quality?
No, not noticeably: 84.78 versus 86.32 in FP16 on the thinking average, i.e. 98.2%. That's practically the level of a 4-bit UD (98.7%) and far ahead of a classic 2-bit quant (84.1%). The loss is measurable, but barely visible in real-world use.
Does it run on Mac?
Yes, via the official MLX variant (Python/Swift, CUDA backend included). Near-lossless 4-bit KV cache, a 1.95 GB DSpark drafter for speculative decoding, an optional 0.63 GB HQQ vision tower. A 32 GB Mac holds the whole thing, including the full 262K context.
Compatible with Ollama or LM Studio?
Not yet for the PTQ1_0/PQ2_0 packings: stock llama.cpp rejects them, and may load a PQ2_0 as a Q2_0 while producing garbage. You need the PrismML-Eng fork. The upstream merge and adoption by mainstream front-ends shouldn't take long, but nothing is scheduled.
How much does it all cost?
€0 on the software side: Apache 2.0 weights on HuggingFace, open source kernels in the fork. The real cost is the hardware — an 8 GB card or a 32 GB Mac — and electricity. To expose the model behind a web interface, hosting from ~€3/month (October 2026, Hostinger) is enough.
I don't have a GPU, what's the alternative?
Three options: the fork's CPU backend (usable, but slow), an Apple Silicon Mac with unified memory, or the free cloud via OpenRouter and Groq — our guide details the settings so you don't sacrifice quality. The 27B class in 2-bit also makes the CPU less ridiculous for occasional inference.
✅ Conclusion
In just one week, 2-bit went from technical curiosity to practical standard: 98.2% of FP16 quality in 5.95 GB — that's the new performance/footprint ratio to beat. Grab the GGUFs from HuggingFace, install the PrismML-Eng fork — and to see where extreme compression leads, our deep dive on 1-bit LLMs charts what's next.