News report
GitHub Copilot Plans HydraFusion Local/Cloud Model Routing for October Preview
GitHub Copilot will test automatic routing between local and cloud models later in October. Here's what works now, what is planned, and the privacy limits.
On this page
GitHub Copilot will test routing coding tasks between local and cloud models
Microsoft and GitHub announced on October 7, 2026 that GitHub Copilot will extend its HydraFusion model orchestrator to models running on a developer's Windows PC. An experimental preview is planned for later in October in the GitHub Copilot app, GitHub Copilot CLI, and Visual Studio Code. This is an announced preview, not a feature that is already generally available.
HydraFusion already routes tasks among cloud models. The announced extension adds a second decision: whether a coding task should use on-device inference or a cloud model. Microsoft says the goal is to balance responsiveness, capability, and inference cost without making developers choose a provider for every turn.
| Capability | Status on October 8, 2026 | Practical boundary |
|---|---|---|
| Copilot CLI local-model discovery | Available from Copilot CLI 1.0.94-0 | The /model picker can discover supported models from a running local Ollama instance; Ollama and the model must already be installed. |
| Manual local model selection | Existing bring-your-own-provider support; newly simplified CLI discovery | Selected models must support tool calling and streaming; the developer confirms which endpoint/model to use. |
| HydraFusion automatic local/cloud routing | Experimental preview planned for later October | Not generally available; supported devices, workloads, policies, and real-world results need verification when the preview ships. |
| Local model execution versus offline mode | Separate controls | Using a local model does not itself disable telemetry or other network activity; Copilot CLI has an explicit offline mode. |
Auto mode considers task context and cached work
The technical announcement describes two paths. In Auto mode, Copilot can consider the task and the state of cached context when deciding whether to use local or cloud inference across a multi-turn coding session. Alternatively, a developer can explicitly select a model or endpoint rather than delegating the choice to the router.
GitHub and Microsoft name MAI Code 1.1 Flash through the Windows ML provider as one local option, alongside models exposed through OpenAI-compatible local endpoints. The announcement emphasizes NVIDIA RTX Spark Windows PCs as a target for powerful local coding, but it does not establish that every Windows PC will run the same large models at useful speed.
The local model still needs enough memory for a real coding session
Microsoft describes MAI Code 1.1 Flash as a mixture-of-experts model with 137 billion total parameters and 6.8 billion active parameters, and says a 3-bit version supports a 256K context window locally. Those are vendor specifications and capability claims, not independent measurements of code quality, tokens per second, or practical memory consumption.
Model weights are only part of the memory requirement. The operating system, other applications, inference runtime, and attention KV cache all consume capacity; long tool-using sessions can increase context size and memory pressure. A large unified-memory configuration may help a model fit, but it is not proof of consistently fast inference.
Sandboxing and model routing solve different problems
The October 7 announcement also discusses GitHub Copilot's generally available local sandboxing through Microsoft Execution Containers. Sandboxing limits what agent-launched tools can access; local/cloud model routing determines where inference runs. Neither feature automatically provides the other.
GitHub documents that shell commands and local MCP servers can run within supported sandbox boundaries when sandboxing is enabled, while built-in file tools use checks in the Copilot harness and remote MCP servers remain outside the local process sandbox. Developers should inspect the effective tool permissions separately from the selected model.
What to watch when the preview arrives
The key questions are which hardware and model providers the experimental preview supports, when Auto chooses local inference, whether switching preserves useful context, and how latency, memory pressure, and cloud-token usage behave during full coding tasks. Microsoft has not provided an independently verified general savings figure for HydraFusion's planned local/cloud routing.
For developers who want to experiment today, the narrower available step is Copilot CLI's local Ollama model discovery in version 1.0.94-0. The promised automatic HydraFusion placement remains a later-October preview, so it should not be represented as a shipping production feature.
Sources
Primary and technical sources
These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.
01 Microsoft Command Line / GitHub
Bringing local models and sandboxed tools to Windows and GitHub Copilot (October 7, 2026)02 Microsoft Windows Experience Blog
Building Windows for hybrid intelligence (October 7, 2026)03 GitHub Changelog
Discover local models in GitHub Copilot CLI (October 7, 2026)04 GitHub Changelog
Copilot CLI now supports BYOK and local models (April 7, 2026)