Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Developers have successfully run the 2.8-trillion-parameter Kimi K3 AI model on consumer Apple Silicon hardware using a custom storage solution called ARGODRIVE. The project prioritizes maintaining full model quality by streaming data from multiple SSDs, achieving a speed of approximately 1 token per second.
Why it matters
This demonstrates that massive, high-parameter AI models previously restricted to enterprise-grade data centers can be executed on consumer hardware, challenging existing assumptions about infrastructure requirements for large language models.
A fork of gavamedia/deltafin (MIT) running Kimi K3 from SSDs on Apple Silicon, with the ARGODRIVE storage work. The benchmark package, placement manifests and results are in k3-public-bench/ ; the measurement instruments are published separately as ARGODRIVE . Credits and what this fork changes: CREDITS.md . The upstream README follows.
____ _ _ __ _ | _ \ ___| | |_ __ _ / _(_)_ __ | | | |/ _ \ | __/ _` | |_| | '_ \ | |_| | __/ | || (_| | _| | | | | |____/ \___|_|\__\__,_|_| |_|_| |_| Run the full, never-pruned , 2.8-trillion-parameter Kimi K3 on consumer hardware, as "fast" as possible Deltafin is a single native binary that runs full Kimi K3. Nothing pruned. Nothing skipped. K3 decides every token.
All 16 experts, every single token. No shortcuts, no "close enough." It's exactly what Moonshot shipped.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in