Skip to content
r/LocalLLaMA

Qwen3.8 27B on RX 7900 XTX: Ollama ROCm vs llama.cpp Vulkan benchmarks

Qwen3.8 27B on RX 7900 XTX: Ollama ROCm vs llama.cpp Vulkan results

A user benchmarked Qwen3.8 27B Q4_K_M on an RX 7900 XTX. Plain decode speed: llama.cpp Vulkan is only ~4% faster than Ollama ROCm (36 vs 34.4 t/s), while Ollama leads in prompt processing at 64K context (215.8 vs 192 t/s). The real gain comes from MTP (multi-token prediction): average generation jumps from 36 t/s to ~69-70 t/s, peaking above 80 t/s. This explains why community reports of 50-80+ t/s are mostly from speculative decoding, not raw single-token decode. The post does not disclose MTP + ngram combination results.

Read the original ↗Export Markdown