Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Developers have created WASTE, an inference engine that allows a 2.78 trillion parameter model to run on a consumer laptop. By streaming experts from disk rather than loading the entire model into RAM, it achieves a functional, albeit slow, performance of 0.5 tokens per second.
Why it matters
This demonstrates that massive, trillion-parameter AI models can be executed on consumer hardware, potentially democratizing access to high-end AI without requiring expensive server-grade infrastructure.
Kimi K3 — 2.78 trillion parameters — running on a consumer laptop.
$ waste run ~/models/k3.waste 'What is the capital of Italy?' waste: no --budget, using 46.24 GB of 64.00 GB (expert cache 17.56 GB) The capital of Italy is **Rome**. [16 tokens, 31.09 s, 0.51 tok/s | experts 3357 hit / 20195 miss = 14%] WASTE is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and uses the remaining RAM as a bounded expert cache.
Its current proof point is the complete open-weights Kimi K3 model: 2.78 trillion parameters, converted into a 982 GiB container and running on a 64 GB MacBook Pro at 0.49–0.54 tokens per second. This is not a distilled, pruned, or reduced variant .
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in