Skip to content
r/LocalLLaMA

Qwen3.6 27B on Dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k Context Working

Qwen3.6 27B on dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k context working

A user ran Qwen3.6 27B with vLLM on dual RTX 5060 Ti 16GB cards, reaching ~62–66 tok/s at 8K. The setup used 32GB VRAM, TP=2, fp8 KV cache, MTP 3 tokens, and a 204800 context window. The tight part is memory: after a 168k prefill, each GPU used ~15.65GiB with max_num_seqs=1.

Why it matters: HKR-H/K/R all pass: the post gives a concrete local-inference benchmark with hardware, vLLM settings, speed, and context limits. Single Reddit sourcing caps it below the 78–84 band.

Read the original ↗Export Markdown