Skip to content
r/LocalLLaMA

32x AMD MI50 32GB runs Kimi K2.6 at 9.7 t/s TG and 264 t/s PP

Final Monster: 32x AMD MI50 32GB at 9.7 t/s (TG) & 264 t/s (PP) with Kimi K2.6

Reddit user ai-infos ran Kimi K2.6 int4 on 32 AMD MI50 32GB GPUs, reaching 9.7 tok/s TG on 136 output tokens. PP hit 263 tok/s on 14,564 input tokens using vllm-gfx906-mobydick across two 16-GPU nodes over 10G Ethernet. Power was about 640W idle and 4,800W peak inference; PCIe bandwidth and the vLLM distributed stack are the real bottlenecks.

Why it matters: HKR-H/K/R all pass via an unusual 32x MI50 build with concrete throughput, power, and network conditions. It stays in the 72–77 band because it is a niche Reddit benchmark, not a broader product or model release.

Read the original ↗Export Markdown