Skip to content
r/LocalLLaMA

Benchmarks of 20 Small LLMs on a 6GB RTX 4050

Benchmarks of 20 small LLMs on a 6GB RTX 4050

The author benchmarked 20 small LLMs on a 6GB RTX 4050 using LM Studio’s OpenAI-compatible API, with N=5 speed runs at 1k, 8k, and 32k context; unsloth/lfm2.5-vl-1.6b led throughput at 207 tok/s on 1k context while using 3.0GB VRAM.

Why it matters: HKR-H/K/R all pass: the low-VRAM GPU hook is concrete, the post gives speed/context/VRAM numbers, and it speaks to local-inference cost pressure. Source authority is a Reddit post, so it stays in the lower featured band.

Read the original ↗Export Markdown