Skip to content
r/LocalLLaMA

2X tk/s on 1× MI50: Qwen3.6-27B inference rises from 19.4 to 38.1 tk/s

2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.

bigattichouse raised Qwen3.6-27B throughput on a single MI50 from 19.4 to 38.1 tk/s by running same-model parallel computations for Q8-or-lower quantization, exploiting unused compute lanes instead of adding a smaller speculative decoding model.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person experiment with numbers and a hypothesis, not a validated release. No code or broader replication is disclosed, so it stays at the featured threshold.

Read the original ↗Export Markdown