Skip to content
r/LocalLLaMA

Qwen 3.8 27B hits 50 tok/s with 100k context on a 16GB GPU

Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)

A user squeezed Qwen 3.8 27B into a 16GB RTX 4070 Ti SUPER, hitting 47–50 tok/s with a 100k-token context window. The trick is beellama.cpp's asymmetric kvarn KV cache quantization—kvarn5 for K, kvarn4 for V—which saved ~6% VRAM and pushed context from 88k to 100k. The model is jrell's IQ4_XS hybrid quant, purpose-built for MTP and long contexts on 16GB cards. MTP speculative decoding with 2 draft tokens and a 1,024-token high-precision tail helped keep quality up. Worth noting: this is a single-user tuning report, not a benchmark, so your mileage may vary.

Why it matters: A community tip with concrete parameters and real-world numbers, directly useful for 16GB GPU owners. Not scored higher because it's a user-shared trick rather than an official release, and beellama.cpp is a niche engine with a smaller audience.

Read the original ↗Export Markdown