News report
Windows ML Adds Experimental llama.cpp and GGUF Support for Local AI
Microsoft's October 7 Windows ML preview adds GGUF via llama.cpp and new text and speech APIs, but production use and C# support remain unavailable.
On this page
- Microsoft brings GGUF models into the Windows ML stack
- Text generation uses a local model and can expose an OpenAI-compatible endpoint
- The preview has important decoding and compatibility limitations
- The native Runtime API adds explicit placement and multi-model pipelines
- PyTorch and Triton updates extend the Windows-on-Arm development path
- What developers should verify before adopting the preview
Microsoft brings GGUF models into the Windows ML stack
Microsoft announced on October 7, 2026 that Windows ML now offers experimental llama.cpp integration for running GGUF language models locally. The same update introduces a preview Windows-native Runtime API and task-specific interfaces for text generation and speech recognition. These are developer-facing preview capabilities, not a general Windows feature that automatically installs a chatbot or makes every model run on every PC.
The practical change is that a developer can bring a GGUF model into Windows ML without first converting it to ONNX. Windows ML selects the llama.cpp backend for GGUF and ONNX Runtime for ONNX/ORT artifacts, while the application remains responsible for choosing and supplying its model. The announcement is distinct from GitHub Copilot's planned HydraFusion local/cloud routing: Windows ML is an inference framework, not a Copilot model-selection service.
| Path | Model/backend | What is supported or limited |
|---|---|---|
| Text Generation API | GGUF via llama.cpp; ONNX/ORT via ONNX Runtime | Experimental; accepts developer-supplied language models and supports streaming and cancellation. |
| Speech Recognition API | ONNX Whisper encoder/decoder plus tokenizer | Experimental; transcribes developer-supplied audio and model, not arbitrary GGUF speech models. |
| Windows ML Runtime API | ONNX Runtime or llama.cpp chosen by artifact format | Experimental native pipelines with explicit CPU/GPU/NPU placement where the backend supports it. |
| GGUF hardware acceleration | llama.cpp CPU module and optional NVIDIA CUDA module | Microsoft validates CPU and CUDA backends for this release; GGUF NPU execution is not supported. |
Text generation uses a local model and can expose an OpenAI-compatible endpoint
The new Text Generation API manages tokenization, autoregressive decoding, streaming output and cancellation over a model the application provides. Microsoft also demonstrates a local Windows ML server exposing an OpenAI-compatible endpoint, allowing existing OpenAI SDK clients to target an on-device model through a loopback address rather than an external model provider.
This compatibility is an API surface, not a promise that a hosted OpenAI model runs locally. Developers still need an appropriate GGUF or ONNX model, sufficient system memory, compatible hardware and the preview runtime. Windows does not bundle a general-purpose language model for this API. A local endpoint can keep inference on the machine, but developers must separately review application telemetry, network access and the source of model files.
The preview has important decoding and compatibility limitations
Microsoft's Text Generation API documentation currently lists greedy decoding as supported but not sampling controls such as top-p, top-k or temperature. It also does not yet expose speculative decoding, multi-token prediction, custom logits processing, chat templates or structured output through that high-level interface. A llama.cpp backend may implement some of these techniques internally, but that does not make them available through the Windows ML Text Generation API.
The backend boundary also matters for hardware claims. Microsoft says its GGUF/llama.cpp integration currently validates CPU execution and an optional CUDA path for NVIDIA GPUs, with the CUDA module deployed alongside the application. It does not establish GGUF inference through every AMD or Intel GPU, nor through an NPU. ONNX/ORT models use the separate ONNX Runtime execution-provider system and may have different hardware support.
The native Runtime API adds explicit placement and multi-model pipelines
For applications needing more control than a text-generation helper, the experimental Windows ML Runtime API can load model artifacts, create CPU, GPU or NPU execution targets where supported, bind tensors, compose processing and model stages, and prepare pipelines for execution. Microsoft's examples include keeping image, audio and video data in Windows-native forms and placing stages on specific devices.
The existing ONNX Runtime API remains supported and is the safer compatibility choice for cross-platform applications or software that cannot accept experimental Windows-only interfaces. The preview native API should be evaluated on each intended model and device rather than treated as an automatic performance upgrade.
PyTorch and Triton updates extend the Windows-on-Arm development path
Alongside Windows ML, Microsoft highlighted official native Windows Arm64 CPU builds of PyTorch, NVIDIA's CUDA-enabled Windows Arm64 PyTorch packages for supported hardware, and Triton work for GPU kernels on Windows. That expands the tooling available for experimentation, training and model preparation on emerging Arm-based Windows systems, including the new RTX Spark class.
These framework changes are not interchangeable with the GGUF runtime announcement. PyTorch and Triton serve development and compilation workflows; Windows ML's task and runtime APIs serve application inference. Microsoft's illustrative performance examples are not independent benchmarks, and real deployment decisions still depend on model compatibility, driver versions and measured behavior on the target machine.
What developers should verify before adopting the preview
The immediate test is narrow: select a supported GGUF model, confirm the Windows ML preview installation and backend, run a local text-generation session, and measure latency and memory usage on the actual target hardware. For speech, the preview requires an ONNX Whisper encoder/decoder model and tokenizer. For production software, Microsoft's current documentation says to keep using supported interfaces rather than ship an experimental Runtime or task API.
Future checkpoints are broader decoding controls, C# bindings, more validated GGUF hardware backends and a stable production support policy. Until Microsoft announces those changes, the confirmed news is new experimental access to GGUF and native inference workflows—not universal model compatibility or a proven speed advantage.
Sources
Primary and technical sources
These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.
01 Microsoft Foundry on Windows Blog
AI Development on Windows: from PyTorch and llama.cpp to Windows ML (October 7, 2026)02 Microsoft Learn
Generate text with your own language model using Windows ML (updated October 7, 2026)03 Microsoft Learn
Windows ML Runtime API overview (updated October 7, 2026)04 Microsoft Learn
Recognize speech with your own Whisper model using Windows ML (updated October 7, 2026)
Related
Continue from here
Useful next steps selected from the same technical reference and publication system.
Technical guide
How to Use System Restore in Windows 11
Use System Restore in Windows 11 from the desktop or Windows Recovery Environment, create restore points, check affected programs, and understand what restoration changes.
Technical guide
Windows Page File Explained: Virtual Memory and Commit Limit
Understand what the Windows page file does, how it extends the system commit limit, how paging differs from RAM use, and why crash dumps can depend on it.
Tool
DDR Memory Latency Calculator
Convert DDR data rate and CAS latency cycles into CAS timing in nanoseconds.
Tool
DDR Memory Bandwidth Calculator
Calculate theoretical peak DDR memory bandwidth from transfer rate, bus width per channel, and active channel count.