Skip to content
Hacker News front page

PrismML's Bonsai 2 27B uses ternary weights to compress a 27B model to 5.9GB while keeping 98.2% of benchmark scores

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

PrismML open-sourced Ternary Bonsai 2 27B, a quantized version of Qwen3.8 27B that uses {-1, 0, +1} weights with FP16 group-wise scaling, hitting 1.76 bits per weight and a 5.9GB footprint — over 9x smaller than the original. It retains 98.2% of the full-precision model's aggregate benchmark score (83.9 vs 85.4), with particularly strong retention in coding, agentic tool use, and vision. Throughput reaches 143 tok/s on an RTX 5090 and 46.8 tok/s on M5 Max; on an RTX 4090 it draws 0.714 mWh/token, 40% more efficient than a full-precision 8B model. The model supports a 262K-token context window, multimodal input, and ships under Apache 2.0. The post does not disclose training data or the specific quantization distillation recipe.

Read the original ↗Export Markdown