Skip to content
r/LocalLLaMA

Qwen3.8-27B GSQ-RCO quant hits highest quality score on llm-bench.io, runs on a single 16GB GPU

qwen3.8-27b-gsq-rco scored higher than any other model yet and fits in 16 GB of VRAM with decent context window — 31.7 tok/s — llm-bench.io

The qwen3.8-27b-gsq-rco quant scored 88.84/100 on llm-bench.io, the highest among 1,100+ community benchmarks. It runs on a single AMD RX 9070 with 16GB VRAM, using 38k of a 96k context window at 31.7 tok/s. The poster says GSQ-RCO preserves quality unusually well and is fast enough for coding. Some commenters call the post and site AI slop, but the model itself gets decent word-of-mouth—one user ran the IQ3_S quant as a daily driver.

Why it matters: Community benchmark #1 + runs on consumer hardware hits all three HKR axes. But source is a Reddit post and third-party benchmark site, not an official release — authority discount keeps it at the featured threshold of 72.

Read the original ↗Export Markdown