Skip to content
AI HOT (Curated Pool)

Swiftlet runs 80B Qwen on Mac with 4.3 GB RAM, 35B on iPhone

Swiftlet:在 Mac 上运行 80B 版 Qwen(内存 4.3 GB),在 iPhone 上运行 35B 版

Swiftlet rewrites MoE inference in Swift + Metal, keeping only a small dense core in memory and streaming expert weights from storage on demand. An 80B Qwen3-Next runs on Mac with 4.3 GB RAM, and a 35B model runs on iPhone. The post doesn't disclose latency or tokens per second, so I'd hold off on real-time expectations.

Why it matters: Swiftlet rewrites MoE inference in Swift + Metal, letting an 80B model run on a Mac with only 4.3 GB and a 35B model on an iPhone — the engineering path is concrete and the numbers are striking, hitting all three HKR axes. The post doesn't give latency or tokens/sec, so real-t...

Read the original ↗Export Markdown