Skip to content
r/LocalLLaMA

Follow-up: Qwen3.6-27B on 1× RTX 3090 reaches ~218K context and stable tool calls

Follow-up: Qwen3.6-27B on 1× RTX 3090 — pushing to ~218K context + ~50–66 TPS, tool calls now stable (PN12 fix)

A Reddit user ran Qwen3.6-27B on one RTX 3090, reporting ~218K context at 50/66 TPS. After fixing Genesis PN12 patch anchor drift, ~25K-token tool outputs stopped OOMing; 198K plus vision reached 51/68 TPS. Single-prompt single-GPU runs still hit a second memory cliff near 50–60K.

Why it matters: HKR-H/K/R all pass: the single-3090 context claim is catchy, the post gives measured TPS and OOM conditions, and local-inference cost pressure resonates. Reddit source keeps it in the low featured band.

Read the original ↗Export Markdown