Qwen3.8-Flash-Next hits 1M context on Strix Halo: 38 tok/s decode, 18 min prefill
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)
Someone ran Qwen3.8-Flash-Next on Strix Halo with halogen 0.12.0 at 1M context. Decode hits 38 tok/s, but prefill takes 18 minutes—too slow for interactive use. The post doesn't specify hardware details or memory bandwidth, but hitting 1M context is a milestone; prefill latency needs work before deployment.