Skip to content
r/LocalLLaMA

Qwen3.8-27B uncensored runs 156K context on one RTX 5090 at 140-190 tok/s

Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2

A Reddit user reports running Qwen3.8-27B uncensored (Q6_K quant) on a single RTX 5090 with 156K context and 140-190 tok/s using DFlash2. The post body is blocked, so no details on inference framework, VRAM usage, or accuracy trade-offs are available.

Read the original ↗Export Markdown