News analysis
AMD Shows Ryzen AI Max+ PRO 495 Local AI Benchmarks Ahead of RTX Spark
AMD's new Gorgon Halo benchmarks put the Ryzen AI Max+ PRO 495 up against Intel in ComfyUI and show what 192GB unified memory can enable locally.
On this page
AMD is putting Gorgon Halo's local-AI pitch into numbers
AMD has shared fresh local-AI performance claims for the Ryzen AI Max+ PRO 495, its 16-core Gorgon Halo flagship with Radeon 8065S graphics and support for up to 192GB of unified memory. The timing puts AMD's high-memory x86 platform back in focus as NVIDIA's RTX Spark Windows PCs approach the same broad local-AI workstation category.
The results are vendor benchmarks, not independent Core Tech Tips measurements. AMD used ComfyUI and compared a 192GB Ryzen AI Max+ PRO 495 system with an Intel Core Ultra X9 388H system configured with 64GB of memory. That unequal memory configuration matters, especially for models that may not fit or run efficiently on the smaller pool.
The largest multiplier needs context
AMD reports a ComfyUI advantage ranging from roughly 1.1x to 32.2x across tested workloads. The extreme result should not be read as evidence that Gorgon Halo is generally thirty times faster than Panther Lake. Tom's Hardware notes that the outlying Yuve workload could be affected by optimization or model-fit limits on the 64GB Intel system.
AMD also reports up to 20 tokens per second for GLM 5.3 Flash using a mixed 4-bit quantization and up to 42 tokens per second for Qwen 3.8 Flash Next using dynamic 4-bit quantization and multi-token prediction. Those are configuration-specific vendor results, not universal LLM throughput figures.
192GB capacity is the bigger architectural story
AMD officially specifies the Ryzen AI Max+ PRO 495 with 16 Zen 5 cores and 32 threads, Radeon 8065S integrated graphics with 40 RDNA 3.5 compute units, an XDNA 2 NPU rated at up to 55 TOPS, and support for as much as 192GB of LPDDR5X memory. AMD says up to 160GB can be allocated as graphics memory.
That capacity can let quantized models fit locally that exceed the practical VRAM ceiling of conventional consumer graphics cards. It does not make unified LPDDR5X equivalent to high-bandwidth discrete GPU memory: model fit, memory bandwidth, compute throughput, quantization, context length and software kernels remain separate constraints.
Sources
Primary and technical sources
These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.