Hacker News·3 min read·medium

Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs

A
Argonautlabs
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
AI Summary

Developers have successfully run the 2.8-trillion-parameter Kimi K3 AI model on consumer Apple Silicon hardware using a custom storage solution called ARGODRIVE. The project prioritizes maintaining full model quality by streaming data from multiple SSDs, achieving a speed of approximately 1 token per second.

Why it matters

This demonstrates that massive, high-parameter AI models previously restricted to enterprise-grade data centers can be executed on consumer hardware, challenging existing assumptions about infrastructure requirements for large language models.

Dive DeeperCreate a free account to unlock

A fork of gavamedia/deltafin (MIT) running Kimi K3 from SSDs on Apple Silicon, with the ARGODRIVE storage work. The benchmark package, placement manifests and results are in k3-public-bench/ ; the measurement instruments are published separately as ARGODRIVE . Credits and what this fork changes: CREDITS.md . The upstream README follows.

____ _ _ __ _ | _ \ ___| | |_ __ _ / _(_)_ __ | | | |/ _ \ | __/ _` | |_| | '_ \ | |_| | __/ | || (_| | _| | | | | |____/ \___|_|\__\__,_|_| |_|_| |_| Run the full, never-pruned , 2.8-trillion-parameter Kimi K3 on consumer hardware, as "fast" as possible Deltafin is a single native binary that runs full Kimi K3. Nothing pruned. Nothing skipped. K3 decides every token.

All 16 experts, every single token. No shortcuts, no "close enough." It's exactly what Moonshot shipped.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in