Hacker News·14 min read
42x faster prompt lookup drafting in llama.cpp
P
pptadversary
[ homepage ] [ github ] [ twitter ] This article was originally published on 2026-09-26.
TL;DR I make drafting for prompt lookup decoding in llama.cpp up to 42x faster while using up to 2.6x less memory through a set of simple performance optimizations largely based on the work of Daniel Lemire and Martin Ankerl .
Update: Daniel Lemire sent in a PR that makes prompt lookup drafting upto 4.2x faster on top of my original optimizations. His work makes the overall speedup upto 140x. I discuss more below .
✦
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in