Swiftlet runs 80B Qwen on Mac with 4.3 GB RAM, 35B on iPhone
Swiftlet:在 Mac 上运行 80B 版 Qwen(内存 4.3 GB),在 iPhone 上运行 35B 版
Swiftlet rewrites MoE inference in Swift + Metal, keeping only a small dense core in memory and streaming expert weights from storage on demand. An 80B Qwen3-Next runs on Mac with 4.3 GB RAM, and a 35B model runs on iPhone. The post doesn't disclose latency or tokens per second, so I'd hold off on real-time expectations.
Why it matters: Swiftlet rewrites MoE inference in Swift + Metal, letting an 80B model run on a Mac with only 4.3 GB and a 35B model on an iPhone — the engineering path is concrete and the numbers are striking, hitting all three HKR axes. The post doesn't give latency or tokens/sec, so real-t...