Kimi K3 Architecture Overview and Notes

An overview of the Kimi K3 architecture, highlighting its transition to a scaled-up version of the Kimi Linear model. The model incorporates efficiency-focused components like LatentMoE and NoPE (No Positional Embeddings) to improve inference performance.
Why it matters
The release of large open-weight models with novel architectural tweaks significantly impacts the competitive landscape of AI development.
The Kimi K3 architecture figure for yesterday’s big open-weight model release, along with some observations and thoughts.
Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
The one new component compared to Kimi Linear is the LatentMoE . I omitted it in the figure below since it’s already very crowded, but that’s essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention .
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in