Qwen3.8 27B at 256K context on a 24GB GPU hits 50 tok/s with MTP
The author runs Qwen3.8 27B at its full 256K context on a single 24GB RTX PRO 4000 SFF, averaging 50.44 tok/s. The gain comes from a custom NVFP4 quant that protects sensitive layers, embedded MTP speculative decoding, and tuned CUDA kernels—not from any single component. Target-only decoding hits 21.19 tok/s; MTP pushes it to 59.46 tok/s. At a nearly full 256K cache, throughput drops to 12.61 tok/s without OOM. The post doesn't disclose total cost, but it's a detailed engineering log, not a plug-and-play recipe.
Why it matters: A solid hands-on local inference post with reproducible numbers and a clear technical path. Hits all three HKR axes, but it's a personal experiment, not an official release or industry event, so it lands at the featured threshold of 78.