Qwen3.8-Flash-Next 177B quantized model achieves 9-10 tokens per second via SSD streaming
What happened
Reddit 用户发帖说,Qwen3.8-Flash-Next 这个 125B 参数的大模型,在 2021 款 32GB 内存的 M1 Max 上跑出了每秒 12-15 个 token 的速度。对消费级硬件来说,这个速度挺亮眼,但帖子正文被 Reddit 屏蔽了,没透露用了什么量化精度、推理框架或优化手段。实际好不好用,还得看精度和上下文长度——这两点...
From r/LocalLLaMA
Coverage
Follow the reports to see the story from different sides.
- r/LocalLLaMAQwen3.8-Flash-Next 177B runs at 9-10 tok/s on a 16 GB GPU via SSD streaming
Qwen3.8-Flash-Next 177B, quantized to NVFP4 (119GiB), achieves 9-10 tok/s on a 16 GB RTX 5060 Ti with 32 GB RAM via SSD streaming. The post is blocked by Reddit and does not disclose implementation details or benchmarks.
- r/LocalLLaMAQwen3.8-Flash-Next 125B runs at 12-15 tok/s on a 2021 M1 Max with 32GB
A Reddit user reports Qwen3.8-Flash-Next 125B runs at 12-15 tok/s on a 2021 M1 Max with 32GB RAM. The post body is blocked by Reddit, so it doesn't disclose quantization, inference framework, or optimization details. The speed is impressive for a 125B model on consumer hardware, but real-world usability depends on precision and context length—neither is specified.
Heat over time
Not enough continuous observations to draw a trend yet.