Skip to content
r/LocalLLaMA

Dual RTX Pro setup hits 150 tok/s decode with Qwen and DeepSeek

Upgraded my local setup with 2 rtx pros and it's amazing.

A developer built a local inference rig with two RTX Pro GPUs running Qwen 3.8 Flash Next and DeepSeek V4 Flash. Decode hits 150 tok/s, prefill 10K tok/s, with room for 4 concurrent requests. He previously used an M3 Ultra 512GB and found it too slow for inference. The new setup handles DeepSeek's 1M context window, but Qwen's thinking tokens eat VRAM. Build took 2.5 days due to a PSU wiring fault. The post doesn't specify GPU model or total VRAM.

Read the original ↗Export Markdown