Wafer runs Kimi K3 on AMD MI355X with better performance per dollar than B300
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Wafer deployed the 2.8T-parameter Kimi K3 on 8 AMD MI355X GPUs, hitting 952 tok/s per node and 48 tok/s per GPU-hour dollar—6.8× the B200 and 1.45× the B300 on perf/dollar. Kimi K3 won't fit on a single 8×B200 node, forcing a slower two-node setup. The MI355X's 288GB HBM per GPU fits the model plus a 1M-token KV cache, giving AMD a practical edge for the first time. The team fixed two ROCm issues: a missing top-k renorm function that broke speculative decoding, and zero-padding attention heads from 12 to 16 to unlock the fast AITER MLA kernel, cutting cold prefill from 51s to 23s. The post doesn't disclose Kimi K3's release date or pricing.
Why it matters: Wafer's real-world benchmark of Kimi K3 on 8× MI355X shows 45% better perf/$ than B300, with concrete numbers and clear methodology. Not scoring higher because it's a single-vendor benchmark with no third-party reproduction, and Wafer sells AMD compute — discount for self-inte...