Skip to content
Hacker News front page

LLMs Are Complicated Now

Ian Barber compares Llama 3 and Nemotron 3 Ultra architectures, showing modern LLMs now pack multiple attention variants, MoE routing, multimodal encoders, and multi-GPU inference. The pattern mirrors how recsys moved from clean two-tower models to complex engineering. The core tension: you can't afford to test a new attention variant without at least partial kernel fusion, but hand-fusing every candidate is too expensive. His takeaway is to design for composability and verifiability upfront, like PyTorch's FlexAttention, so the research loop stays cheap.

Why it matters: A concrete architecture comparison that puts Llama 3 and Nemotron 3 Ultra diagrams side by side, walking through complexity growth across attention variants, routing, multimodal encoders, and inference deployment. Hits all three HKR axes, but it's observational synthesis rathe...

Read the original ↗Export Markdown