Show HN: Shoehorn – Quantize any model down to run on your machine

Shoehorn is a new tool that allows users to quantize and run large language models locally based on their specific hardware memory constraints. It optimizes model performance by calculating the exact memory budget available on the user's machine.
Why it matters
This tool democratizes access to powerful AI models by allowing users with limited hardware to run high-quality models locally without needing expensive enterprise-grade GPUs.
Make any language model fit the memory you actually have.
Preset quantizations ignore your hardware: pick one that fits and you either waste hundreds of megabytes of quality headroom or find out at load time it didn't fit after all. shoehorn starts from the memory you actually have, subtracts what inference itself needs, and solves a per-tensor mixed-precision assignment that lands within a rounding error of the remainder — routinely using 99.99% of the budget, sometimes to the byte.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in