Running Kimi K3 on a M1 Mac
This technical guide explains how to run the Kimi K3 large language model on an Apple M1 Mac using a research project called Deltafin. It details the hardware requirements and software commands needed to manage the model's massive parameter size through local storage or streaming.
Why it matters
It demonstrates the feasibility of running massive AI models on consumer-grade hardware, pushing the boundaries of local machine learning capabilities.
____ _ _ __ _ | _ \ ___| | |_ __ _ / _(_)_ __ | | | |/ _ \ | __/ _` | |_| | '_ \ | |_| | __/ | || (_| | _| | | | | |____/ \___|_|\__\__,_|_| |_|_| |_| An experiment in running Kimi K3 (2.8T parameters) on one Apple Silicon Mac Deltafin is a small research project that runs a Mixture-of-Experts model far larger than the machine it sits on. It is not fast — about 16 seconds per token on our M1 Max — but it is exact, reproducible, and it works on a 64 GB laptop. Newer chips and more RAM make it faster automatically.
Three commands, then you're generating. The only real decision is step 3.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in