Skip to content
r/LocalLLaMA

Qwen3.8-Flash-Next 125B runs at 12-15 tok/s on a 2021 M1 Max with 32GB

Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max

A Reddit user reports Qwen3.8-Flash-Next 125B runs at 12-15 tok/s on a 2021 M1 Max with 32GB RAM. The post body is blocked by Reddit, so it doesn't disclose quantization, inference framework, or optimization details. The speed is impressive for a 125B model on consumer hardware, but real-world usability depends on precision and context length—neither is specified.

Read the original ↗Export Markdown