Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
imec's aistack team benchmarked 64 real coding tasks across self-hosted GPUs, rented hardware, and commercial APIs. The newly added Kimi K3, running on an 8×B300 node, costs about 20% more in hardware than the 8×B200 setup used for GLM-5.2, but hits 86.4% task resolution—roughly 24 points above both GLM-5.2 and Claude Opus 4.8 at 62.5%. The trade-off is speed: K3 handles 16 concurrent sessions with a median task time of 38 minutes, about 8× slower than the Claude Code baseline. The post flags that SWEBench Pro tasks may have leaked into K3's training data, so take that resolution number with a grain of salt. The core takeaway: self-hosting doesn't save money—you buy hardware for peak load but pay for it 24/7, and utilization is what makes or breaks the cost case.
Why it matters: imec benchmarked self-hosted, rented, and API setups on 64 real coding tasks. Kimi K3 on 8×B300 hit 86.4% completion—nearly 24pp above GLM-5.2 and Claude Opus 4.8—at a 20% hardware premium. Solid data with named models and numbers; direct signal for teams deciding on self-host...