Microsoft is taking its AI push out of cloud computing and is launching a new strategy in which more AI processing is directly on Windows PCs.
The company has also developed what it calls “Windows for hybrid intelligence” in which AI tasks can be done locally on a computer when needed and the higher-value work is sent to cloud-based models instead.
The announcement was made by Microsoft Executive Vice President for Windows and Devices Pavan Davuluri on October 7. Microsoft CEO Satya Nadella also made it clear in San Francisco that the strategy was key to the company’s new Windows and Surface hardware.
The approach might allow users and developers to run certain AI workloads without always making use of cloud infrastructure. Microsoft says local processing can offer benefits in responsiveness, privacy and the economics of running AI workloads.
MAI Code 1.1 Flash is For Local PCs
Microsoft's core announcement is a local version of MAI Code 1.1 Flash, a coding-based artificial intelligence model that runs directly on compatible Windows PCs.
Microsoft says the model has 137 billion total parameters, with 6.8 billion active parameters. The company has used quantisation and speculative decoding to reduce the model footprint and improve responsiveness while maintaining the capabilities needed for coding and agentic tasks.
The local model uses 53GB of memory, according to Microsoft's announcement. It also supports a 256K token context window so it can deal with a lot of information while coding.
Microsoft will also support even more local models for Windows. These include an upcoming NVIDIA Nemotron model with more than 70 billion parameters and DeepSeek V4 Flash on systems based on NVIDIA RTX Spark hardware. Windows ML also will support llama.cpp, giving developers access to open source models and local AI experiments.
But running these models locally requires a lot of powerful hardware and substantial memory. That means that the technology first needs to be made to work for high-performance PCs as opposed to your regular laptops and desktops.
GitHub Copilot Will Decide Between Local And Cloud AI
Microsoft is also adding hybrid to GitHub Copilot.
The coding assistant will be able to tell if a particular task is better done with an on-device model or cloud-based model. The goal is to balance those factors (speed, performance and computing costs) without manually deciding where each task should run.
Microsoft said the capability will be available by the end of October and will be available first for NVIDIA RTX Spark Windows PCs, including the new Surface Laptop Ultra.
That’s how a developer could apply local AI to a coding task that is relevant and send more challenging jobs to cloud-scale models. Microsoft believes this hybrid setup will provide users with better control over performance, cost and privacy.
New Security Tools for AI Agents.
Since AI agents can access files, run commands and interact with applications, security is one of the most prevalent concerns.
Microsoft has therefore introduced Microsoft Execution Containers (MXC), a technology to limit what an AI agent can access and do on a computer.
MXC is now available for Windows and provides a containment layer around agent workloads. Developers can add resources such as files and network destinations that an agent needs, and the system enforces those boundaries. An AI agent or its generated code cannot simply grant itself further access.
Microsoft says the system is built around three main areas: containment, identity and manageability. Containment limits an agent's access, identity helps distinguish actions taken by an agent from actions taken by a person and manageability allows organisations to govern and monitor agent activity.
In addition, several AI tools and frameworks are already supporting MXC, such as GitHub Copilot, OpenAI Codex, OpenClaw, Replit, LM Studio and Unsloth AI. Anthropic's Claude Code and many other platforms are also expected to support it.
Microsoft Bets On The Hybrid AI PC.
Microsoft’s latest strategy is part of a larger transition in how AI would work on personal computers.
Instead of viewing the PC as just a device that connects users to cloud AI services, Microsoft wants Windows machines to combine local computing with cloud intelligence. Microsoft calls this hybrid intelligence.
The company will also launch new hardware for these workloads. The Surface Laptop Ultra, with NVIDIA RTX Spark, is a high-performance Windows machine for demanding AI, development and creative workloads. Microsoft will be taking pre-orders for that device with the Surface RTX Spark Dev Box.
Microsoft’s approach could also reduce reliance on cloud processing for workloads that can be handled locally. Users should also have access to cloud models when tasks require more computing power.
That puts Microsoft in the middle of a growing competition to make personal computers capable of running AI. Apple is also pursuing on-device AI while chipmakers like NVIDIA are developing hardware for the increasingly demanding local models.
The recent Microsoft announcements indicate that the next AI computing phase will not be about choosing between the PC and the cloud. But the most productive AI systems will know when to use both.