Skip to content
r/LocalLLaMA

Qwen 3.6 wins benchmarks, but Gemma 4 looks stronger in local vision tests

Qwen 3.6 wins the benchmarks, but Gemma 4 wins reality. 7 things I learned testing 27B/31B Vision models locally (vLLM / FP8) side by side. Benchmaxing seems real.

A Reddit user compared Qwen 3.6 and Gemma 4 locally on vLLM FP8 across 27B/31B vision models. Qwen burned 8,000+ tokens on hard GeoGuessr cases, while Gemma often used 1,500; Qwen also needed 2 FPS video preprocessing. The practitioner detail: vLLM and Llama.cpp can default Gemma visual tokens to 280, while 1,120+ improved fine-detail accuracy.

Why it matters: HKR-H/K/R all pass: the post has a sharp benchmark-vs-reality hook and concrete local vLLM/FP8 settings. A single Reddit test limits authority, so it sits just above the featured threshold.

Read the original ↗Export Markdown