Skip to content
Hacker News front page

Slotstream runs the 104GB Qwen3.8-Flash-Next on a 48GB Mac at ~12 tok/s

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

carloslfu open-sourced Slotstream, an MLX + Swift tool that runs the 125B-parameter MoE model Qwen3.8-Flash-Next (104GB at 4-bit) on a Mac with only 48GB RAM. It streams expert modules from SSD on demand instead of loading everything into memory. Speed is ~12 tok/s, and it exposes an Ollama-compatible API. The post doesn't disclose time-to-first-token or the SSD model used, so real-world feel is still an open question.

Why it matters: Streams MoE experts from SSD on demand via MLX + Swift, letting a 104GB Qwen 125B model hit ~12 tok/s on a 48GB Mac. Clean engineering with an Ollama-compatible API that lowers the trial barrier. Docked a few points because it's a solo project with no community validation or m...

Read the original ↗Export Markdown