I clicked on this because Ian Barber nailed a pain point anyone doing model R&D hits daily. He puts Llama 3 and Nemotron 3 Ultra architecture diagrams side by side, and you can see at a glance how much stuff is packed into modern LLMs: grouped-query attention, compressed attention, sparse attention, sliding windows, MoE routing, and now selective routing on residual streams and attention blocks too. Vision and audio encoders went from bolt-on to baked-in, and inference runs across multiple GPUs with comms ops spliced into the middle of your model.
The recsys comparison is spot-on. Recommendation systems went from clean two-tower models to engineering monsters for the same reason: you need more capability without killing inference efficiency. But LLMs have it worse. If you want to try a new attention variant and it's an order of magnitude slower than the current one, you can't even tell if it's worth pursuing. The current variant is fused and optimized, so your candidate needs at least partial fusion just to produce comparable performance numbers. That's the trap: hand-fusing is too slow, and auto-generating kernels doesn't give you a trustworthy baseline to verify correctness.
His answer is PyTorch's FlexAttention, which uses Triton templates to generate kernels for a whole class of attention ops, designed upfront for composability and verifiability. I buy this, but the scope is limited—FlexAttention handles attention, and you've still got MoE routing, multimodal encoders, and cross-GPU comms that all need the same treatment, with no equivalent tooling yet.
He mentions Andrej Karpathy joining Anthropic to work on auto-research loops, arguing that cutting architectures to their essence and making them composable matters as much as building clever agent setups. Interesting take in mid-2026 context, but the post doesn't elaborate, so treat it as a footnote.