Running Qwen3.6-35B-A3B on a laptop RTX 4060 8GB
A Reddit user ran Qwen3.6-35B-A3B on an RTX 4060 8GB laptop and reported that --no-mmap raised generation from about 11 to 43 tok/s, while speculative decoding with a Qwen3.5-0.8B draft model improved throughput by 26%.
Why it matters: HKR-H/K/R all pass: the post has a clear laptop-35B hook, reproducible speed numbers, and local-LLM resonance. Reddit single-post sourcing keeps it below the 78+ good-quality band.