Microsoft pushes local code: MAI-Code-1.1 Flash (137B) runs in 3-bit on a Windows PC
🔎 The day frontier code left the datacenter
On October 8, 2026, Microsoft crossed a line no major vendor had yet dared to cross so head-on: publishing a 3-bit build of its frontier code model, MAI-Code-1.1 Flash, designed to run locally on a high-end Windows PC. Not a lab prototype, not a research paper — a downloadable model, wired into GitHub Copilot, benchmarks included (Microsoft).
The numbers are dizzying: 137 billion total parameters, 6.8 billion active per token thanks to the mixture-of-experts architecture, quantization at roughly 3.3 bits per weight, and 75.5 GB of memory at peak — with a full 256k-token context. Eighteen months ago, this was pure fantasy. Today, it's a spec sheet.
And Microsoft doesn't stop at the model. GitHub Copilot becomes a local/cloud router, Windows Hybrid Intelligence orchestrates it all, and Execution Containers sandbox what the model generates. This isn't a product announcement. It's a declaration of war on 100% cloud inference.
The essentials
- The model: MAI-Code-1.1 Flash in a 3-bit build (~3.3 bits per weight), a 137B-parameter MoE (6.8B active), llama.cpp CUDA runtime on Windows ARM64.
- The setup: more than 120 GB of RAM recommended (128 GB for best results), 75.5 GB peak memory at 256k context.
- The performance: 923.5 tokens/s in prompt processing at 64k, 769.8 tokens/s at 128k.
- The quality: 70.8% on SWE-Bench Verified in the quantized on-device version, versus 72.6% at full precision — barely a 1.8-point loss.
- The integration: GitHub Copilot delegates to the local model or the cloud (Auto mode), with experimental support in the app, the CLI, and VS Code by the end of October.
- The entry price: the Surface Laptop Ultra 128 GB, the official target machine, at $2,899 (October 2026, check on microsoft.com).
Recommended tools
| Tool | Main use | Price (October 2026) | Ideal for |
|---|---|---|---|
| MAI-Code-1.1 Flash (3-bit build) | Local 137B code model, 256K context | Free (download) | Developers with 128 GB machines |
| GitHub Copilot | Code agent + local/cloud routing | ≈$10/month (Copilot Pro, check on github.com) | Devs who want the best of both worlds |
| Ollama | Local LLM runtime in CLI | Free | Automating and scripting inference |
| LM Studio | GUI for local LLMs | Free | Discovering local AI without the command line |
| llama.cpp | Low-level runtime, total control | Free | Squeezing out every token/s |
| Hostinger | GPU VPS to offload inference | from a few $/month (check on hostinger.com) | Those who don't have 128 GB of RAM |
What Microsoft announced on October 8, with the numbers to back it up
A downloadable 3-bit build of MAI-Code-1.1 Flash, optimized to run locally on Windows — with performance that calls into question the systematic reliance on the cloud. The announcement, notably relayed by ExplainX, marks a turning point in Microsoft's AI strategy.
The model itself is not new. Microsoft launched it on August 11, 2026 with a simple promise: better code quality at a quarter of the cost of version 1.0, 25% fewer tokens per task, and 25% faster streaming (Microsoft AI). At the time, Microsoft was claiming +22% on Terminal-Bench 2.1 in Copilot CLI and +15% on .NET tasks. The model already runs in production in GitHub Copilot on the cloud side.
What changes in October is the distribution. The local build compresses the weights to ~3.3 bits, embeds DFlash2 speculative decoding, and runs via llama.cpp with CUDA on Windows ARM64. The official model card confirms a 256K token context and text + image multimodal input (model card, PDF).
A note of journalistic honesty: Microsoft's technical post describes 137B parameters (6.8B active), whereas the model card states 138B (5B active). The discrepancy probably comes from a different way of counting embedding layers. The order of magnitude, however, is not in dispute.
The benchmarks: quantification barely breaks anything
| Benchmark | MAI-Code-1.1 Flash (full) | Local 3-bit build | GPT-OSS-120B |
|---|---|---|---|
| SWE-Bench Verified | 72.6% | 70.8% | 32.0% |
| Terminal-Bench 2.1 | 62.9% | 66.29% | 23.6% |
Source: Microsoft, October 2026.
Two readings stand out. First, the loss due to quantization is marginal: 1.8 points on SWE-Bench Verified between the full version and the on-device build. Second — and this is more surprising — the 3-bit version performs better than the full version on Terminal-Bench 2.1 (66.29% vs 62.9%). Aggressive compression can sometimes smooth out certain agentic behaviors in a beneficial way.
The gap with GPT-OSS-120B, the mainstream open-source reference cited by Microsoft, is abyssal. The message is crystal clear: even compressed to 3-bit, the Redmond model crushes what the community could run on consumer hardware until now.
128 GB of RAM: the bill that hurts
Yes, you really do need more than 120 GB of memory — it's the entry price for running a frontier model at home. Microsoft recommends more than 120 GB of RAM, with 128 GB for the best results (Neowin, October 7, 2026).
The technical details explain this requirement. At full context (256k tokens), the model peaks at 75.5 GB of memory. Add the OS, your IDE, GitHub Copilot and the rest, and you understand why Microsoft promises nothing below 120 GB. WindowsReport confirms that the 3-bit version still retains the full 256K context (WindowsReport) — Microsoft did not sacrifice the context window on the altar of compression.
Throughput-wise, the numbers are respectable for a model of this size: 923.5 tokens/s in prompt processing at 64k context, 769.8 tokens/s at 128k. In other words, feeding the agent a large codebase remains smooth — that's often where local models used to fall short.
The target machine Microsoft openly embraces? The Surface Laptop Ultra in top configuration, 128 GB, priced at $2,899 (October 2026). While marketing departments sell "AI PCs" with 16 GB of memory, Microsoft has just set the real bar. We're a long way from "local AI for everyone" — here, we're talking about a machine costing nearly $3,000.
No 128 GB on hand? Three plan Bs
- Offload the inference: a GPU VPS rented on demand, for example from Hostinger, lets you run the model without changing machines.
- Aim smaller: the best local LLMs like Qwen3.6-27B or DeepSeek V4 Flash run on far more modest setups.
- Wait for what's next: the recent history of compression suggests the hardware bar will drop quickly. More on that below.
GitHub Copilot as conductor: local/cloud routing
Copilot isn't becoming a local model — it's becoming the router that decides, task by task, what stays on your machine and what goes to the cloud. This is perhaps the most strategic part of the announcement.
Concretely, GitHub Copilot adds local model selection in two modes: Auto orchestration, where the system distributes the work between local and cloud on its own, or explicit model selection via Windows ML or any OpenAI-compatible endpoint. Support will be experimental by the end of October in the Copilot app, the CLI, and VS Code (Neowin).
Microsoft's most compelling demonstration: everything works in airplane mode. Disconnected computer, local model, coding agent. For companies that prohibit sending proprietary code to third-party APIs, it's a killer selling point — and a market that cloud pure players struggle to address.
This orchestration is a direct continuation of Build 2026: MAI-Thinking, the Copilot Super App, and the first proprietary models. The logic is the same: Windows stops being a mere OS and becomes a hybrid AI platform, where local/cloud routing is a system primitive.
Execution Containers: sandboxing goes GA
Running a model that writes and executes code locally raises an obvious question: how do you stop an agent from doing whatever it wants on your machine? Microsoft's answer: Microsoft Execution Containers (MXC), now available in stable release, with a ProcessContainer BaseContainer backend. On macOS, isolation is handled by Seatbelt; on Linux, by bubblewrap.
It's a detail that says a lot. Microsoft isn't just building a local model — it's building the security rails around it. That's exactly what most homebrewed local inference stacks are missing, where the agent often runs with permissions that are far too broad.
The technique: 3-bit, MoE, and speculative decoding
Three building blocks make the feat possible — and none of them is magic. This is compression engineering applied methodically, not a theoretical breakthrough.
Quantization at ~3.3 bits per weight. A conventional model stores each weight in 16 bits. At 3.3 bits, the memory footprint is divided by five. The community had shown the way: our feature on the 2-bit ternary revolution, Bonsai 2, and the 27B GGUFs already detailed how these techniques run multimodal models on consumer GPUs. Microsoft is industrializing the approach, with published benchmarks to prove that quality keeps up.
Mixture-of-experts with 6.8B active parameters. Out of the 137B total parameters, only 6.8B are activated per token. The result: memory holds the entire model, but per-token computation remains that of a model roughly twenty times smaller. It's the best of both worlds — the knowledge capacity of a giant, the speed of a mid-sized model.
DFlash2 speculative decoding. A small "draft" model proposes several tokens ahead, and the large model validates them in parallel. On hardware where memory bandwidth is the bottleneck — and on a PC, it always is — this is what makes it possible to maintain usable generation speeds despite the model's size.
The combination of these three building blocks explains the announced figures: a 137B model that fits in ~75 GB and swallows 128k-token prompts at nearly 800 tokens/s. Two years ago, that required an H100 node costing several dollars per hour.
Windows vs Mac: the local inference war opens
Microsoft is not alone in this space — and that's precisely what makes the announcement interesting. On Apple's side, Salvatore Sanfilippo (antirez, the creator of Redis) has launched ds4, a local inference engine that makes DeepSeek V4 Flash usable on a Mac. Two ecosystems, two philosophies, one shared conclusion: local has become a battlefield.
The comparison is illuminating. Apple is betting on the unified memory of its chips and on community engines like ds4 to run open-weights models. Microsoft is betting on Windows machines with 128 GB of RAM, a runtime based on llama.cpp, and its own proprietary model. In both cases, the thesis is identical: the best model in the world is useless if it forces your data to leave your machine.
The strategic difference is notable. Apple lets the ecosystem build; Microsoft builds it itself and locks down the experience end to end, from silicon to sandbox. To dig deeper into the Mac path, our article on ds4 and DeepSeek V4 Flash on Mac details antirez's approach.
Where does MAI-Code-1.1 Flash stand against the competition?
In the "local code" niche, MAI-Code-1.1 Flash has no real competitor — but it doesn't yet rival the top of the cloud. 72.6% on SWE-Bench Verified is solid and well above GPT-OSS-120B (32.0%), but below what the proprietary heavyweights promise.
To situate the landscape, here's our ranking of the best LLMs for coding:
| Model | Score | Access | Positioning |
|---|---|---|---|
| GPT-5.5 (OpenAI) | 98.2 | Cloud | The absolute peak, but everything goes through remote servers |
| Gemini 3 Pro Deep Think (Google) | 95.4 | Cloud | Heavy reasoning for complex tasks |
| Claude Opus 4.7 Adaptive (Anthropic) | 94.3 | Cloud | The benchmark for coding agents |
| DeepSeek V4 Pro (Max) (DeepSeek) | 88 | Open weights | Open source frontier |
| Kimi K2.6 (Moonshot AI) | 85 | Open weights | Versatile, agentic-oriented |
| GLM-5.1 (Z.AI) | 83 | Open weights | Excellent performance-to-resource ratio |
| DeepSeek V4 Flash (DeepSeek) | 76 | Open weights | The local star, via ds4 on Mac |
| Qwen3.6-27B (Alibaba) | 74 | Open weights | Runs on consumer hardware |
Warning: these aggregate scores are not directly comparable to Microsoft's SWE-Bench Verified — the scales differ. But the hierarchy is telling: MAI-Code-1.1 Flash running locally sits in the solid open-source pack, not in the cloud top three.
And that's exactly the right positioning. The local model doesn't need to beat GPT-5.5. It has to be good enough for the bulk of everyday tasks — refactoring, tests, documentation, small features — while keeping the code on the machine. The cloud, meanwhile, keeps the gnarly problems.
Analysis: frontier local inference goes mainstream
The real signal in this announcement isn't the model — it's the distribution. When the world's largest software vendor pushes a frontier model into consumer PCs, with orchestration and sandboxing included, local inference stops being an enthusiast hobby and becomes an industrial option.
First consequence: privacy becomes a selling point, not a niche. Microsoft's airplane mode isn't a demo gimmick. It's a direct answer to the CIOs who block sending source code to third-party APIs. An enterprise developer is worth far more than a consumer developer in value terms — and that's precisely the profile that local inference unlocks.
Second consequence: the model's economics change logic. Recall the August 2026 figures: a quarter of the cost of version 1.0, 25% fewer tokens per task (Microsoft AI). Microsoft is optimizing cloud cost and local feasibility simultaneously. That's consistent: the more efficient the model, the less memory it demands, and the wider the pool of compatible machines grows.
Third consequence: the developer's job shifts, it doesn't disappear. Those who fear that local AI will make programming obsolete are missing the point: these tools demand a better understanding of code, not ignoring it. Our article on no-code vs code: when should you learn to program? is still relevant — the more powerful the tools, the more the person who understands what they're doing gains the edge over the one who delegates blindly.
The limitations, to be honest: 128 GB of RAM is elite hardware; Copilot support is still experimental; and 72.6% on SWE-Bench Verified doesn't rival GPT-5.5. But trajectories matter more than fixed points — and the trajectory is not in doubt.
How to try MAI-Code-1.1 Flash locally
The model is downloadable now, and GitHub Copilot support arrives in experimental form by the end of October 2026. Here's the shortest path.
- Check your setup. More than 120 GB of RAM (128 GB ideal), Windows ARM64 and a CUDA-compatible GPU. Below that, skip it or aim for a smaller model.
- Download the 3-bit build from the official Microsoft page.
- Test via llama.cpp if you want to control every parameter — the official runtime is built on it:
llama-server --model MAI-Code-1.1-Flash-3bit.gguf --ctx-size 65536
- Hook up Copilot as soon as experimental support lands in the app, the CLI and VS Code: choose Auto mode (local/cloud) or force the local model via Windows ML.
- Sandbox it. Enable Execution Containers for everything the model generates and runs — it's GA, there's no excuse.
If you're new to local AI, our guide to installing a local LLM with Ollama or LM Studio remains the best starting point: the logic is identical, only the model size changes. And to keep up with upcoming announcements of this kind, the AI news radar is updated continuously.
❌ Common Mistakes
Mistake 1: Believing 75 GB of RAM is enough
The 75.5 GB figure corresponds to peak memory usage at full context (256k), model only. Add the OS, your IDE, Copilot, and a safety margin: Microsoft recommends more than 120 GB. Solution: aim for 128 GB, or deliberately reduce the context if your tasks allow it.
Mistake 2: Trying to replace the cloud with local
At 70.8% on SWE-Bench Verified, the local build is excellent for everyday use, but it can't compete with GPT-5.5 (98.2 on our index). Solution: use Copilot's Auto mode, which routes tricky tasks to the cloud and keeps local for the bulk of the work.
Mistake 3: Running generated code without a sandbox
A local model that writes and runs code on your machine without any isolation is an open door to disaster. Solution: Execution Containers (MXC) are now available as a stable release — use them, or their native equivalents (Seatbelt on macOS, bubblewrap on Linux).
Mistake 4: Believing "3-bit" means "cut-rate quality"
Quantization at ~3.3 bits per weight only costs 1.8 points on SWE-Bench Verified — and even gains points on Terminal-Bench 2.1. Solution: judge by the published benchmarks, not by the bit count.
❓ Frequently Asked Questions
What PC do you need to run MAI-Code-1.1 Flash locally?
More than 120 GB of RAM (128 GB recommended), Windows ARM64, and a CUDA-compatible GPU. At the full 256k token context, the model alone climbs to 75.5 GB of memory. The Surface Laptop Ultra 128 GB ($2,899, October 2026) is the reference machine cited by Microsoft.
Is the 3-bit build free?
The model is downloadable for free from Microsoft's official pages. The full orchestration experience, however, goes through GitHub Copilot and its subscription (≈$10/month for Copilot Pro, October 2026, check github.com). The llama.cpp runtime, for its part, remains open source.
Can you run it on a Mac?
The official build targets Windows ARM64 with CUDA. On Mac, the credible alternative goes through ds4, antirez's engine, which makes DeepSeek V4 Flash usable on Apple Silicon. Same thesis — the code stays on the machine — different ecosystems.
Does the local model replace my Copilot subscription?
No. Copilot orchestrates the routing between local and cloud, and the hardest tasks are better off going to GPT-5.5 or Claude Opus 4.7. The local build covers the everyday — refactoring, tests, documentation — while the cloud covers the exceptional. The two are complementary.
Does MAI-Code-1.1 Flash understand images?
Yes. The official model card indicates text + image input with a 256K token context. Enough to send interface screenshots, mockups, or architecture diagrams directly to the agent, without depending on an external OCR tool.
When is support in GitHub Copilot coming?
Experimental support by the end of October 2026 in the Copilot app, the CLI, and VS Code, according to Neowin. On the cloud side, MAI-Code-1.1 Flash has already been in production in Copilot since its launch on August 11, 2026.
✅ Conclusion
By releasing a 3-bit build of MAI-Code-1.1 Flash with Copilot orchestration and GA sandboxing, Microsoft has just turned local inference of frontier models from a technical exploit into a product feature. To choose your next coding machine, our comparison of the best AI tools for coding is up to date.