Skip to content
Hacker News front page

Ternary LLMs break the 1.58-bit floor: BITCOS hits 1.485 bits per weight by exploiting zero-weight density

Breaking the 1.58-bit Barrier for Ternary LLMs

Ternary models store weights as -1, 0, or +1, with a theoretical floor of ~1.585 bits and a practical 1.625 bits in five-trit packing. Intel authors measured 29 ternary LLMs and found up to 51.5% zeros. BITCOS replaces fixed packing with a presence bitmap plus a compacted sign vector, costing 2 minus zero-density bits per weight. It beats five-trit packing on 26 of 29 models and reaches 1.485 bits on the sparsest. Optimized unpacking on AVX-512, AVX2, and Xe2 GPUs yields up to 1.28× faster matrix-vector multiply; end-to-end decode throughput improves up to 1.18× on CPUs and 1.27× on GPUs. The paper does not name the models or disclose their parameter counts, nor whether they are publicly available.

Why it matters: Intel team measured 29 ternary models, found zero weights up to 51.5%, and proposed BITCOS encoding to break the 1.58-bit floor. Solid K with concrete numbers and a new mechanism; H works on title intrigue. But it's a narrow inference-opt topic with no R pull, so it lands at t...

Read the original ↗Export Markdown