Article may be outdated

This article is 66 days old. Some details may have changed since publication.

Hacker News·5 min read·hard

You Could Have Come Up with Kimi Delta Attention

A
AnhTho_FR
You Could Have Come Up with Kimi Delta Attention
✦AI Summary

This technical article explains the mathematical derivation of Kimi Delta Attention, a variant of linear attention mechanisms. It uses bra-ket notation to simplify the understanding of complex state update equations in modern AI models.

Why it matters

Understanding these efficient attention mechanisms is crucial for developers working on optimizing large language model inference and architecture.

✦Dive DeeperCreate a free account to unlock

You Could Have Come Up With Kimi Delta Attention Jamie Dborin Founder & Member of Technical Staff, Doubleword Math notation ⟨k|q⟩ kᵀq A note on notation: this article defaults to bra-ket notation because (in my quantum-inspired opinion) it makes the shapes in this derivation very clear. The Math notation switch above rewrites every equation using conventional bold vectors and explicit transposes instead. In bra-ket mode, ∣ q ⟩ \lvert q\rangle ∣ q ⟩ is a column vector, ⟨ k ∣ \langle k\rvert ⟨ k ∣ is a row vector, ⟨ k ∣ q ⟩ \langle k\rvert q\rangle ⟨ k ∣ q ⟩ is a number, and ∣ v ⟩ ⟨ k ∣ \lvert v\rangle\langle k\rvert ∣ v ⟩ ⟨ k ∣ is a matrix. Vectors face right by default, while keys face left when written into the linear-attention state.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in