News analysis

Vulkan Beats ROCm on Llama.cpp Throughput in Fresh AMD Linux Tests

Fresh Lemonade/Llama.cpp tests on Strix Halo and Radeon AI PRO R9700 show Vulkan usually leading throughput while ROCm usually cuts time to first token.

On this page
  1. Fresh AMD tests split the win between throughput and latency
  2. Strix Halo favored Vulkan for sustained generation
  3. The discrete Radeon AI PRO R9700 showed the same broad pattern

Fresh AMD tests split the win between throughput and latency

Fresh independent Linux testing from Phoronix shows that choosing between the Vulkan and ROCm back ends for llama.cpp is not a simple one-winner decision. Using Lemonade 2026.39.1 with llama.cpp b10825, Vulkan generally delivered higher token-generation throughput, while AMD's ROCm path generally produced a lower time to first token.

The pattern appeared on two substantially different AMD systems: a Framework Desktop with Ryzen AI Max+ 395 Strix Halo integrated Radeon graphics, and a System76 Thelio Major workstation with a discrete Radeon AI PRO R9700. That cross-platform consistency makes the result more useful than a single-machine comparison, but it still belongs to this exact software stack, models and test scenarios.

What the October 5 Lemonade tests found
MetricVulkanROCm
Token throughputUsually faster across the tested models and scenariosUsually behind Vulkan, with some exceptions
Time to first tokenUsually slower on the tested AMD systemsUsually lower latency; sometimes by a large margin
Hardware coverage in this testRyzen AI Max+ 395 and Radeon AI PRO R9700Ryzen AI Max+ 395 and Radeon AI PRO R9700
RuntimeLemonade 2026.39.1 / llama.cpp b10825Lemonade 2026.39.1 / llama.cpp b10825

Strix Halo favored Vulkan for sustained generation

On the Ryzen AI Max+ 395 Framework Desktop, Vulkan repeatedly produced more tokens per second in Qwen3 14B, Qwen3.5 4B, MiniCPM4 8B and several other tested scenarios. ROCm, however, repeatedly reached the first generated token sooner. Phoronix also recorded exceptions: ROCm won throughput in at least one Qwen3-Coder-Next code-debug scenario.

That distinction matters for local AI use. Sustained tokens per second affects how quickly a long response is generated after output begins, while time to first token affects how responsive the model feels before the first visible output. A backend that wins one metric can therefore still lose the other.

The discrete Radeon AI PRO R9700 showed the same broad pattern

The Radeon AI PRO R9700 workstation broadly repeated the Strix Halo result despite using different hardware and a different Ubuntu/kernel stack. Vulkan generally led tokens per second, while ROCm generally delivered lower latency to the first token. There were again workload-level exceptions, including cases where Vulkan also won first-token latency or ROCm won throughput.

This is stronger evidence for a backend-level tendency in the tested llama.cpp/Lemonade configuration, but it is not evidence that Vulkan is universally faster than ROCm for AI. ROCm supports a much broader compute ecosystem than this one llama.cpp comparison, and results can move with model architecture, quantization, runtime version, kernels, drivers and GPU generation.

Sources

Primary and technical sources

These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.

  1. 01 Phoronix

    AMD ROCm vs. Vulkan Performance For Lemonade Local AI Server With Llama.cpp
  2. 02 AMD

    Lemonade Server