Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
Soup is a new command-line tool designed to simplify the fine-tuning of Large Language Models on consumer-grade hardware. It uses layer streaming and quantization to allow users to train 8B models on GPUs with as little as 4 GB of VRAM.
Why it matters
It lowers the barrier to entry for AI development, enabling individual developers to fine-tune powerful models without expensive enterprise infrastructure.
Fine-tune and post-train LLMs in one command. No SSH, no config hell.
Website · Quick Start · Config · Docs · Commands · Models · Discord
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
pip install " soup-cli[train] " # add [train] to fine-tune; bare `soup-cli` is the light CLI soup init --template chat soup train Why Soup? Training LLMs is still painful. Even experienced teams spend 30-50% of their time fighting infrastructure instead of improving models. Soup fixes that.
v0.72.4 — align on a laptop: DPO, ORPO, SimPO and KTO over layer streaming. Layer streaming keeps the frozen base out of VRAM and feeds it to the GPU one decoder layer at a time. It used to support supervised fine-tuning only; now it runs the preference losses too.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in