Skip to content
r/LocalLLaMA

Qwen3.8 27B is great for local agentic coding — what hardware to upgrade to?

Qwen3.8 27b for agentic coding and next .... what?

A user runs Qwen3.8 27B quantized on a single RTX 3090 (24GB VRAM) with 100K context and finds it excellent for agentic coding. They argue small local models can handle 80-90% of mundane coding work without API costs, posing a real threat to closed-source vendors. But upgrading to a smarter model reveals a gap: Kimi-K3 is too large, MiniMax-M3 is too slow on dual 3090s. They ask what hardware others use for frontier-level models and whether multi-GPU servers are worth it. Comments suggest dual RTX 5060 Ti 16GB can run Qwen3.8 27B at 50 t/s, but design phases still rely on frontier models.

Read the original ↗Export Markdown