Skip to content
r/LocalLLaMA

Qwen3.8-Flash-Next 177B runs at 9-10 tok/s on a 16 GB GPU via SSD streaming

Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

Qwen3.8-Flash-Next 177B, quantized to NVFP4 (119GiB), achieves 9-10 tok/s on a 16 GB RTX 5060 Ti with 32 GB RAM via SSD streaming. The post is blocked by Reddit and does not disclose implementation details or benchmarks.

Read the original ↗Export Markdown