Skip to content
Hacker News front page

ByteShape releases full ShapeLearn quant for Qwen 3.8 27B, hitting 99.63% accuracy at 13.1 GB VRAM

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

ByteShape released full ShapeLearn quantized models for Qwen 3.8 27B. All five models sit on the quality-speed frontier across six GPUs. Default pick GPU-5 (IQ4_XS, 3.84 bpw, 13.1 GB) scores 99.63% of BF16 at 93.66 tok/s on an RTX 5090. If VRAM is tight, GPU-4 (IQ3_S, 11.0 GB) still delivers 98.72% accuracy and runs faster. Supports MTP and DFlash2 speculative decoding; DFlash2 is faster for text-only but needs an extra 1.1 GB VRAM and doesn't handle images. The post doesn't disclose training data or optimization budget details.

Read the original ↗Export Markdown