The Reddit body is blocked with a 403, so the usable evidence is the title and summary numbers only. That matters. We do not have quantization details for every run, context length, batch size, prompt length, CUDA version, PCIe topology, offload settings, or flash-attention settings. Even with that caveat, the signal is strong: Mistral Medium 3.5 128B on 4× RTX 3080 20GB moves from 10.37 t/s to 21.59 t/s under llama.cpp tensor split; Qwen 3.5 122B A10B drops from 60.08 t/s to 53.49 t/s; vLLM serves Qwen GPTQ-Int4 at 187.04 tok/s.
My read is simple: this is not a cute LocalLLaMA tuning post. It shows local inference pressure moving away from raw VRAM capacity and toward parallelism policy, memory bandwidth, and MoE routing overhead. Four RTX 3080 20GB cards give 80GB of VRAM, on older Ampere consumer hardware. That is enough to squeeze 120B-class models only through low-bit quantization and careful sharding. The question is no longer just whether the model loads. The question is whether prefill survives, long context stays usable, and decode speed remains stable under real workloads.
The Mistral Medium 3.5 128B result fits the expected shape. A dense 128B model benefits from spreading tensor work across four cards. The reported tg128 speed rises from 10.37 to 21.59 t/s, about a 2.08× gain. That is useful, but it is nowhere near 4×. PCIe communication, KV-cache behavior, synchronization, and kernel scheduling eat the rest. An RTX 3080 20GB has strong local bandwidth, roughly in the high hundreds of GB/s, but consumer multi-GPU boxes lack the clean interconnect story of server rigs. llama.cpp tensor split can reduce per-card pressure. It cannot make cross-GPU movement free. The result looks credible for exactly that reason: a large improvement, but not a fantasy scaling curve.
The Qwen 3.5 122B A10B result is the sharper part. The summary frames it as MoE, and A10B likely means roughly 10B active parameters per token. MoE saves compute by activating a subset of experts. It does not make the system easy. Routing, expert placement, batch shape, and cross-card layout all matter. If llama.cpp tensor split pushes Qwen from 60.08 t/s down to 53.49 t/s, I do not find that surprising. A naive tensor-parallel layout can turn expert access into fragmented cross-card traffic. The active-parameter advantage gets taxed by communication. Dense models often like tensor parallelism. MoE models need a more careful placement story.
The vLLM number is the one most likely to be abused. 187.04 tok/s for Qwen GPTQ-Int4 is not a direct victory over the 53.49 t/s llama.cpp line unless the original post used the same prompt, same batch size, same quantization target, same context length, and same measurement method. The visible body does not disclose that. vLLM is strong at PagedAttention, continuous batching, and server-style scheduling. llama.cpp is strong at broad hardware reach, format flexibility, and low-friction local use. If 187.04 tok/s is aggregate throughput under batching, it says more about the serving path than about single-user interactive latency.
There is useful context here from the last wave of local inference work. The community has moved from “can a 70B fit on consumer hardware?” to “can 100B-plus low-bit models deliver usable throughput?” Qwen’s MoE route echoes the Mixtral 8×7B moment: large total parameter count, much smaller active compute per token, and plenty of system-side pain. Mixtral also exposed how much expert placement, quant format, and batching policy change real throughput in llama.cpp-style stacks. If Qwen 3.5 122B A10B really reaches around 60 t/s on 4×3080s under a reproducible setup, that is enough for many local agent backends, code-edit loops, batch rewriting jobs, and small internal tools. It is not automatically enough for long-context chat with heavy prefill.
I have two doubts. First, the visible article does not show whether the numbers include prefill. tg128 usually tells a decode-heavy story, and decode alone flatters local setups. Feed the system 8K or 32K tokens, and prefill can wreck the user experience. Second, the vLLM 187.04 tok/s figure may be batch throughput rather than single-stream speed. Those are different metrics. Local benchmark posts often put both on one chart, and people then quote the highest number as if it describes every workload.
I would still feature this. Open model deployment boundaries are often discovered in posts like this before they show up in polished vendor material. Mistral Medium 3.5 128B and Qwen 3.5 122B A10B running on 4× RTX 3080 20GB puts 120B-class inference inside the serious hobbyist and small-lab zone. For many teams, that is more operationally relevant than another H100 cluster benchmark. The hardware and model names are disclosed, and the summary gives the speed figures. The missing reproduction details stop this from being a universal claim. The clean takeaway is narrower: dense 120B models can benefit materially from llama.cpp tensor split, MoE models are highly sensitive to parallel layout, and vLLM still has a strong serving-path advantage when the workload matches its scheduler.