42x Faster Prompt Lookup Drafting in llama.cpp

[homepage] [github] [twitter]

    <sub>This article was originally published on 2026-09-26.</sub>

TL;DR I make drafting for prompt lookup decoding in llama.cpp up to 42x faster while using up to 2.6x less memory through a set of simple performance optimizations largely based on the work of Daniel Lemire and Martin Ankerl.

Update: Daniel Lemire sent in a PR that makes prompt lookup drafting upto 4.2x faster on top of my original optimizations. His work makes the overall speedup upto 140x. I discuss more below.

Many popular inference engines including llama.cpp and vllm, and machine learning libraries such as hugging face's transformers library, support prompt lookup decoding (also called n-gram speculation) for faster token generation. Prompt lookup decoding is technically a special case of speculative decoding that uses a really stupid draft model, an n-gram model. When prompt lookup decoding is used, the inference engine drafts the next $k$ tokens using the following rule.