Frontier AI on Your Own Hardware
Tim Dettmers's dlab is open-sourcing a full stack this week to run frontier AI on local hardware. An agent auto-optimized Metal kernels to run Qwen 3.6 35B-A3B at 1.5 bits per weight, hitting 450 tokens/s on a Mac. The core argument: the unit of research is no longer the paper but a coherent ecosystem. Full details are still under wraps, but the release includes an autonomous research agent, efficient test-time scaling, and auto-compaction that beats Claude Code on token savings.
Why it matters: Tim Dettmers is a key figure in quantization, and this isn't a single paper but a full toolchain release with concrete numbers (1.5 bits, 450 tok/s) and a reproducible path. The deduction: it's a blog announcement — actual usability and compatibility won't be clear until the o...