Skip to content
r/LocalLLaMA

Qwen3.8-Flash-Next: 256K context, 16 tok/s on DDR4 + Tesla T4

Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4

Qwen3.8-Flash-Next achieves 16 tok/s on DDR4 RAM and a Tesla T4 GPU with 256K context. This combination shows long-context models can run on low-cost hardware, useful for local deployment and edge scenarios. The post is blocked by Reddit and does not disclose architecture, training method, or release date.

Read the original ↗Export Markdown