Qwen 3.8-27B NVFP4 beats Q5_K_M and nears BF16 after tweaking temp and min_p
Qwen 3.8-27B NVFP4 actually beats Q5_K_M and touches official BF16 levels, but only if you change 2 sampling params. Also almost 3x faster.
A user compared Qwen 3.8-27B NVFP4 (NInfer) against Q5_K_M (llama.cpp) on a 5090. With default sampling, Q5 led on IFBench strict (76% vs 74%), but NVFP4's loose score was 78%, meaning its failures were mostly minor formatting drift. After lowering temp to 0.9 and setting min_p to 0.05, NVFP4 hit 80% strict while Q5 stayed at 76%. The official BF16 baseline is 79.5% on 300 samples; NVFP4's 80% on 50 samples has variance but the trend held across reruns. NVFP4 also ran nearly 3x faster and used far less VRAM. The takeaway: default sampling hides NVFP4's real quality—tighten it and you get near-BF16 instruction following.
Why it matters: First-person benchmark with concrete numbers and a reproducible recipe — not armchair theory. Directly useful for local LLM users. Score capped because it's a single-GPU, single-model case relying on the niche NInfer engine, limiting generalizability.