Introducing BreadBowl-Embed  ·  Open weights  ·  Apache-2.0

One vector is too few. One per token is too many.

BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text.

01 — The problem

High-precision retrieval reads your documents twice.

Once to find them, and again to decide which ones actually matter.

Modern search and RAG pipelines run in two stages. First, a bi-encoder compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance (the date, the exception, who did what to whom), a cross-encoder re-reads the top candidates next to the query and scores them again.