Article may be outdated

This article is 61 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

M
marcobambini
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
✦AI Summary

Developers have created WASTE, an inference engine that allows a 2.78 trillion parameter model to run on a consumer laptop. By streaming experts from disk rather than loading the entire model into RAM, it achieves a functional, albeit slow, performance of 0.5 tokens per second.

Why it matters

This demonstrates that massive, trillion-parameter AI models can be executed on consumer hardware, potentially democratizing access to high-end AI without requiring expensive server-grade infrastructure.

✦Dive DeeperCreate a free account to unlock

Kimi K3 — 2.78 trillion parameters — running on a consumer laptop.

$ waste run ~/models/k3.waste 'What is the capital of Italy?' waste: no --budget, using 46.24 GB of 64.00 GB (expert cache 17.56 GB) The capital of Italy is **Rome**. [16 tokens, 31.09 s, 0.51 tok/s | experts 3357 hit / 20195 miss = 14%] WASTE is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and uses the remaining RAM as a bounded expert cache.

Its current proof point is the complete open-weights Kimi K3 model: 2.78 trillion parameters, converted into a 982 GiB container and running on a 64 GB MacBook Pro at 0.49–0.54 tokens per second. This is not a distilled, pruned, or reduced variant .

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in