AirLLM 70B inference with single 4GB GPU
AirLLM is a software tool that enables the execution of massive large language models on consumer-grade hardware with limited VRAM. By utilizing per-expert streaming for sparse models, it allows users to run models as large as 2.8 trillion parameters on a single GPU.
Why it matters
This technology significantly lowers the barrier to entry for running state-of-the-art AI, democratizing access to powerful models without requiring expensive enterprise-grade hardware.
Quickstart | Configurations | MacOS | Example notebooks | FAQ
AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , DeepSeek-V3 (671B) on ~12GB , and Kimi K3 (2.8T) — the largest open-source model released to date — on under 4GB , because sparse MoE models stream one expert at a time rather than a whole layer.
Bloome — build & run AI agent teams in the cloud, zero setup
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in