Microsoft Brings AI On-Premises, But You'll Need 120GB of RAM
Microsoft unveils its hybrid intelligence strategy: MAI-Code-1.1-Flash can run locally, GitHub Copilot integrates HydraFusion, and MXC reaches general availability. Local inference is free of charge, but requires fairly high hardware specs.
On October 7, Microsoft announced a whole slate of new updates. The core focus isn't any single model, but "hybrid intelligence" — cutting-edge AI doesn't have to live exclusively in the cloud anymore, it can now run locally on your device too.
Let's start with the model: MAI-Code-1.1-Flash. It's an MoE architecture with 137B total parameters and 6.8B active parameters, built specifically for real-world coding scenarios. The version unveiled at Build ran on the cloud; this time Microsoft optimized it for local deployment. With 3-bit quantization, the model size was shrunk by nearly 80% while retaining a 256K context window, and its performance on SWE-Bench Verified and Terminal-Bench 2.1 comes close to matching the full-precision version.
The most interesting part is how it's deployed. Microsoft has built MAI-Code-1.1-Flash directly into GitHub Copilot's routing layer. Previously, every Copilot request was routed through the cloud; now there's a local option: if your machine can handle it, the task runs entirely locally. Only tasks that your hardware can't handle or that require more powerful capabilities are forwarded to the cloud. Calls to the local model do not count against your token billing.
One netizen crunched the numbers: the quantized model weights are around 53GB, peak memory usage hits roughly 75.5GB when running with the full 256K context window, and Microsoft officially recommends 120GB of RAM or more. In other words, saving on token costs requires you to already own a machine with high RAM capacity. As one netizen put it bluntly: "Local inference cuts your token bill, but shifts the cost to electricity bills and hardware wear and tear."

There are several other accompanying updates.
GitHub HydraFusion has been extended to Windows. Originally a cloud-based multi-model routing project, it can now call local models. It will launch as an experimental preview in the Copilot desktop app, CLI, and VS Code by the end of the month.
MXC (Microsoft Execution Containers) has officially reached general availability. It's a sandbox for AI agents that controls which files and networks an agent can access, with policies enforced at runtime. OpenAI Codex, GitHub Copilot, Replit, and LM Studio already support it, with Anthropic Claude Code, Perplexity, Manus and more on the way.
llama.cpp has been integrated into Windows ML. This means more open-source models can run on Windows-powered GPUs, NPUs, and CPUs, significantly lowering the barrier for developers to test out new models.
On the hardware side, Microsoft partnered with NVIDIA for this rollout. The Surface Laptop Ultra featuring the RTX Spark platform comes with 128GB of unified memory to run models with over 120B parameters locally, starting at $2,599. The Surface RTX Spark Dev Box starts at $5,999. For higher-end use cases, there's also DGX Station for Windows, equipped with a GB300 Ultra chip and 748GB of memory, built to target one-trillion-parameter models.
Microsoft named this full-stack offering "hybrid intelligence", and the logic is complete: models, runtime, containers, routing, and hardware are all fully integrated. But one detail stands out — the official 120GB+ RAM recommendation is far beyond what the average consumer uses. Microsoft is targeting developers, AI engineers, and enterprises strained by soaring cloud token bills.
There's also a noteworthy signal: Microsoft says upcoming Copilot updates will enable it to "understand your PC" — local files, recent activity, and device diagnostics can all be used as context. AI is no longer just a cloud-hosted chat box; it's starting to live on your computer.
Whether this offering succeeds depends on two things: how many of these high-RAM devices can Microsoft sell, and whether local model quality can hold up for real-world development workflows. Microsoft has left the choice up to developers: want free tokens? Buy the hardware first.
Pretty shrewd business.
发布时间: 2026-10-08 18:44