Article may be outdated

This article is 17 days old. Some details may have changed since publication.

Hacker News·2 min read·hard

Show HN: Shoehorn – Quantize any model down to run on your machine

R
rhgraysonii
Show HN: Shoehorn – Quantize any model down to run on your machine
AI Summary

Shoehorn is a new tool that allows users to quantize and run large language models locally based on their specific hardware memory constraints. It optimizes model performance by calculating the exact memory budget available on the user's machine.

Why it matters

This tool democratizes access to powerful AI models by allowing users with limited hardware to run high-quality models locally without needing expensive enterprise-grade GPUs.

Dive DeeperCreate a free account to unlock

Make any language model fit the memory you actually have.

Preset quantizations ignore your hardware: pick one that fits and you either waste hundreds of megabytes of quality headroom or find out at load time it didn't fit after all. shoehorn starts from the memory you actually have, subtracts what inference itself needs, and solves a per-tensor mixed-precision assignment that lands within a rounding error of the remainder — routinely using 99.99% of the budget, sometimes to the byte.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technology

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in