News report

Windows ML Adds Experimental llama.cpp and GGUF Support for Local AI

Microsoft's October 7 Windows ML preview adds GGUF via llama.cpp and new text and speech APIs, but production use and C# support remain unavailable.

On this page
  1. Microsoft brings GGUF models into the Windows ML stack
  2. Text generation uses a local model and can expose an OpenAI-compatible endpoint
  3. The preview has important decoding and compatibility limitations
  4. The native Runtime API adds explicit placement and multi-model pipelines
  5. PyTorch and Triton updates extend the Windows-on-Arm development path
  6. What developers should verify before adopting the preview

Microsoft brings GGUF models into the Windows ML stack

Microsoft announced on October 7, 2026 that Windows ML now offers experimental llama.cpp integration for running GGUF language models locally. The same update introduces a preview Windows-native Runtime API and task-specific interfaces for text generation and speech recognition. These are developer-facing preview capabilities, not a general Windows feature that automatically installs a chatbot or makes every model run on every PC.

The practical change is that a developer can bring a GGUF model into Windows ML without first converting it to ONNX. Windows ML selects the llama.cpp backend for GGUF and ONNX Runtime for ONNX/ORT artifacts, while the application remains responsible for choosing and supplying its model. The announcement is distinct from GitHub Copilot's planned HydraFusion local/cloud routing: Windows ML is an inference framework, not a Copilot model-selection service.

New Windows ML developer paths and their current boundaries (October 8, 2026)
PathModel/backendWhat is supported or limited
Text Generation APIGGUF via llama.cpp; ONNX/ORT via ONNX RuntimeExperimental; accepts developer-supplied language models and supports streaming and cancellation.
Speech Recognition APIONNX Whisper encoder/decoder plus tokenizerExperimental; transcribes developer-supplied audio and model, not arbitrary GGUF speech models.
Windows ML Runtime APIONNX Runtime or llama.cpp chosen by artifact formatExperimental native pipelines with explicit CPU/GPU/NPU placement where the backend supports it.
GGUF hardware accelerationllama.cpp CPU module and optional NVIDIA CUDA moduleMicrosoft validates CPU and CUDA backends for this release; GGUF NPU execution is not supported.

Text generation uses a local model and can expose an OpenAI-compatible endpoint

The new Text Generation API manages tokenization, autoregressive decoding, streaming output and cancellation over a model the application provides. Microsoft also demonstrates a local Windows ML server exposing an OpenAI-compatible endpoint, allowing existing OpenAI SDK clients to target an on-device model through a loopback address rather than an external model provider.

This compatibility is an API surface, not a promise that a hosted OpenAI model runs locally. Developers still need an appropriate GGUF or ONNX model, sufficient system memory, compatible hardware and the preview runtime. Windows does not bundle a general-purpose language model for this API. A local endpoint can keep inference on the machine, but developers must separately review application telemetry, network access and the source of model files.

The preview has important decoding and compatibility limitations

Microsoft's Text Generation API documentation currently lists greedy decoding as supported but not sampling controls such as top-p, top-k or temperature. It also does not yet expose speculative decoding, multi-token prediction, custom logits processing, chat templates or structured output through that high-level interface. A llama.cpp backend may implement some of these techniques internally, but that does not make them available through the Windows ML Text Generation API.

The backend boundary also matters for hardware claims. Microsoft says its GGUF/llama.cpp integration currently validates CPU execution and an optional CUDA path for NVIDIA GPUs, with the CUDA module deployed alongside the application. It does not establish GGUF inference through every AMD or Intel GPU, nor through an NPU. ONNX/ORT models use the separate ONNX Runtime execution-provider system and may have different hardware support.

The native Runtime API adds explicit placement and multi-model pipelines

For applications needing more control than a text-generation helper, the experimental Windows ML Runtime API can load model artifacts, create CPU, GPU or NPU execution targets where supported, bind tensors, compose processing and model stages, and prepare pipelines for execution. Microsoft's examples include keeping image, audio and video data in Windows-native forms and placing stages on specific devices.

The existing ONNX Runtime API remains supported and is the safer compatibility choice for cross-platform applications or software that cannot accept experimental Windows-only interfaces. The preview native API should be evaluated on each intended model and device rather than treated as an automatic performance upgrade.

PyTorch and Triton updates extend the Windows-on-Arm development path

Alongside Windows ML, Microsoft highlighted official native Windows Arm64 CPU builds of PyTorch, NVIDIA's CUDA-enabled Windows Arm64 PyTorch packages for supported hardware, and Triton work for GPU kernels on Windows. That expands the tooling available for experimentation, training and model preparation on emerging Arm-based Windows systems, including the new RTX Spark class.

These framework changes are not interchangeable with the GGUF runtime announcement. PyTorch and Triton serve development and compilation workflows; Windows ML's task and runtime APIs serve application inference. Microsoft's illustrative performance examples are not independent benchmarks, and real deployment decisions still depend on model compatibility, driver versions and measured behavior on the target machine.

What developers should verify before adopting the preview

The immediate test is narrow: select a supported GGUF model, confirm the Windows ML preview installation and backend, run a local text-generation session, and measure latency and memory usage on the actual target hardware. For speech, the preview requires an ONNX Whisper encoder/decoder model and tokenizer. For production software, Microsoft's current documentation says to keep using supported interfaces rather than ship an experimental Runtime or task API.

Future checkpoints are broader decoding controls, C# bindings, more validated GGUF hardware backends and a stable production support policy. Until Microsoft announces those changes, the confirmed news is new experimental access to GGUF and native inference workflows—not universal model compatibility or a proven speed advantage.

Sources

Primary and technical sources

These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.

  1. 01 Microsoft Foundry on Windows Blog

    AI Development on Windows: from PyTorch and llama.cpp to Windows ML (October 7, 2026)
  2. 02 Microsoft Learn

    Generate text with your own language model using Windows ML (updated October 7, 2026)
  3. 03 Microsoft Learn

    Windows ML Runtime API overview (updated October 7, 2026)
  4. 04 Microsoft Learn

    Recognize speech with your own Whisper model using Windows ML (updated October 7, 2026)

Related

Technical guide

How to Use System Restore in Windows 11

Use System Restore in Windows 11 from the desktop or Windows Recovery Environment, create restore points, check affected programs, and understand what restoration changes.