The Inference Engineering Masterclass with Baseten's Philip Kiely and Ali Taha
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten just raised a $13B Series F. Philip Kiely and Ali Taha explain why inference engineering is now its own discipline, covering quantization, speculative decoding, KV-cache movement, and disaggregated prefill/decode. In one GLM-5.2 experiment, quantizing more layers preserved benchmark quality while boosting throughput 20% because errors across layers canceled out. They also detail grafting Kimi's vision encoder onto GLM-5.2 without touching the language model, and note that inference optimizations can still deliver 20% to 200% gains. The conversation touches on NVIDIA Dynamo, Rubin, video generation, and local inference, but the post doesn't expand on those.
Why it matters: Baseten's $13B raise gives this deep-dive on inference engineering extra timeliness. The GLM-5.2 quantization experiment and Kimi vision encoder graft are concrete, novel details. Score stays at 78 rather than higher because it's a podcast transcript — high signal density but ...