Skip to content
r/LocalLLaMA

2x RTX 3090 setup for local Qwen 3.6 27B inference

we really all are going to make it, aren't we? 2x3090 setup.

A Reddit user ran Qwen 3.6 27B on a dual RTX 3090 Ubuntu setup, reporting 48GB VRAM, a 262k context window, no NVLink, about 4000 pp/s prompt processing, and 113 tk/s generation.

Why it matters: All HKR axes pass, and this is a first-person local-inference run with concrete numbers. Source is a single Reddit post with limited reproducibility detail, so it sits at the low featured threshold.

Read the original ↗Export Markdown