Shoehorn: Quantize any model to fit your exact memory budget, down to the byte
Show HN: Shoehorn – Quantize any model down to run on your machine
Shoehorn is an open-source quantizer that starts from your available memory, subtracts inference overhead, then solves a per-tensor mixed-precision assignment that routinely uses over 99.99% of the budget. It avoids preset quantization tiers that either waste hundreds of megabytes or fail at load time. The local web UI measures your machine, streams the fit, shows perplexity cost, and launches a chat. It requires llama.cpp on PATH, outputs standard GGUF v3, and runs on macOS Apple Silicon, Linux, and Windows. The quantizer is written from scratch in Rust.
Why it matters: Open-source tool with a genuinely useful inversion of the quantization problem — measure first, allocate later. The 99.99% utilization number is concrete. But it's a solo dev's Show HN project with no paper or large-scale validation, so it stays at the featured threshold of 78.