Skip to content
r/LocalLLaMA

Dual Radeon PRO R9700 hits 111 tok/s on Qwen 3.8 27B, costs less than one RTX 5090

2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next

A hobbyist built a dual Radeon AI PRO R9700 (32 GB each) rig for local inference. With vLLM Radiance and Qwen 3.8 27B MXFP4, median decode hit 111.4 tok/s, ITL 1% low 77.9 tok/s, TTFT 81 ms, and ~7k prefill reached 4,410 tok/s. FP8 was slower at 87.6 tok/s decode. Qwen 3.8 Flash Next with GGUF and expert offload on a SATA SSD managed 35.4 tok/s; the author expects a bump with NVMe. The whole build cost ~€4,000, over €1,000 less than a single 32 GB RTX 5090. The post does not include concurrency sweeps or KV cache degradation data—those are planned next.

Read the original ↗Export Markdown