Muse Glimmer fits an agent on-device with a memory hierarchy disguised as a 30B Transformer
Muse Glimmer is a memory hierarchy disguised as a 30B Transformer
Meta's Muse Glimmer is a ~30B multimodal model built to run agentic tasks offline on consumer hardware. The BF16 checkpoint is 55 GiB; Meta ships ~4-bit quantized versions that bring the language model under 20 GB. Architecturally, only every fourth of the 52 layers uses full-context attention—the other 39 use a 2,048-token sliding window. Global layers drop RoPE and retrieve by content. The KV cache stores just two key/value heads while 32 query heads provide diverse retrieval behaviors. A large ViT handles perception once, compresses neighboring patches 4:1, and feeds them as tokens. The result is a memory hierarchy: local layers build ordered representations, global layers search across the full sequence, and the tiny KV cache means quantization savings translate directly into longer context or larger batches.
Why it matters: A solid architecture deep-dive with real numbers on quantization cost, attention hierarchy, and QK norm. But it's a third-party analysis, not a Meta launch, and the pure-architecture focus raises the bar for readers outside on-device deployment — so it lands right at the featu...