Qwen3.8-Flash-Next hits 84 tok/s on Strix Halo + RTX 3090 Ti, nearly matching dual-3090 server
Qwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time
A Reddit user tested the 104 GB MoE model Qwen3.8-Flash-Next on a Strix Halo laptop with an external RTX 3090 Ti, boosting inference from 22 to 84 tok/s. It scored within one HumanEval+ problem of a dual-3090 vLLM server at 0.4x the wall time. The post is blocked by Reddit and does not disclose hardware details, quantization, or power draw.