24+ tok/s from ~30B MoE models on an old GTX 1080
24+ tok/s from ~30B MoE models on an old GTX 1080 (8 GB VRAM, 128k context)
User mdda ran Qwen 3.6 35B-A3B on an i7-6700, GTX 1080, and 32GB RAM machine at about 24 tok/s with 128k context; the setup uses llama.cpp MoE offloading plus TurboQuant/RotorQuant KV cache quantization, with PCIe 3.0 x16 saturated and GPU utilization at about 40–50%.
Why it matters: Single Reddit source limits authority, but the GTX 1080 + Qwen 3.6 35B-A3B + 128k + 24 tok/s setup gives a concrete local-inference result. HKR-H/K/R all pass; this is a practical featured item, not a major model or product launch.