Skip to content
Hacker News front page

Shoehorn: Quantize any model to fit your exact memory budget, down to the byte

Show HN: Shoehorn – Quantize any model down to run on your machine

Shoehorn is an open-source quantizer that starts from your available memory, subtracts inference overhead, then solves a per-tensor mixed-precision assignment that routinely uses over 99.99% of the budget. It avoids preset quantization tiers that either waste hundreds of megabytes or fail at load time. The local web UI measures your machine, streams the fit, shows perplexity cost, and launches a chat. It requires llama.cpp on PATH, outputs standard GGUF v3, and runs on macOS Apple Silicon, Linux, and Windows. The quantizer is written from scratch in Rust.

Why it matters: Open-source tool with a genuinely useful inversion of the quantization problem — measure first, allocate later. The 99.99% utilization number is concrete. But it's a solo dev's Show HN project with no paper or large-scale validation, so it stays at the featured threshold of 78.

Read the original ↗Export Markdown