Skip to content
r/LocalLLaMA

Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB

A Reddit user benchmarked Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB. llama.cpp tensor split raised Mistral tg128 from 10.37 to 21.59 t/s, but Qwen MoE fell from 60.08 to 53.49 t/s. vLLM served Qwen GPTQ-Int4 at 187.04 tok/s; the key signal is MoE sensitivity to parallel strategy.

Why it matters: HKR-H/K/R all pass: the 4×RTX 3080 setup is a strong hook, and the post gives concrete llama.cpp/vLLM throughput deltas. Reddit single-run sourcing keeps it in the 72–77 band.

Read the original ↗Export Markdown