Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K
Long-context coding on RTX 5080 16GB: Qwen3.6-35B-A3B holds 30 t/s at 128K (89 t/s fresh), no quality drop
A Reddit user tested a local coding-agent setup on RTX 5080 16GB; the title says Qwen3.6-35B-A3B reaches 30 t/s at 128K. The post lists Ryzen 9700X, 96GB DDR5, Windows 11, and CUDA 12.9.1 as required. Qwen3.6-27B dense hit only 3.2 t/s at 128K, so the key path is KV quantization plus MoE offload.
Why it matters: HKR-H/K/R all pass: 30 t/s at 128K on a 16GB RTX 5080 is a strong hook, with hardware/CUDA details and a dense baseline. Single Reddit run lacks multi-source reproduction, so featured not P1.