{"article":{"slug":"one-vector-is-too-few-one-per-token-is-too-many","title":"One vector is too few. One per token is too many.","subtitle":null,"summary":"BreadBowl-Embed introduces a late-interaction embedding architecture that stores each passage as 16 routing–value slots—aiming for cross-encoder-like precision without re-reading documents, between single-vector and per-token extremes.","content_type":"blog_post","language":"en","canonical_url":"https://breadbowl.ai/blog/breadbowl-embed/","author":{"name":"Ming (Jerry) Xu","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"BreadBowl","url":"https://breadbowl.ai","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3524,"reading_minutes":15,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-04T11:14:33.094Z","updated_at":"2026-10-04T11:14:33.094Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/one-vector-is-too-few-one-per-token-is-too-many","markdown_url":"https://listedarticles.com/articles/one-vector-is-too-few-one-per-token-is-too-many.md","example":false,"citation":"Ming (Jerry) Xu, BreadBowl. \"One vector is too few. One per token is too many..\" 2 Oct 2026. https://breadbowl.ai/blog/breadbowl-embed/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://breadbowl.ai/blog/breadbowl-embed/"},"body_markdown":"**Introducing BreadBowl-Embed**  ·  Open weights  ·  Apache-2.0\n\n# One vector is too *few*. One per token is too *many*.\n\n    BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text.\n\n01 — The problem\n\n## High-precision retrieval reads your documents twice.\n\nOnce to find them, and again to decide which ones actually matter.\n\nModern search and RAG pipelines run in two stages. First, a **bi-encoder** compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance (the date, the exception, who did what to whom), a **cross-encoder** re-reads the top candidates next to the query and scores them again.\n\nThat second stage is where the precision comes from, and it is expensive in a very specific way: the reranker runs a full transformer pass over every (query, candidate) pair, every time a query arrives. Most of that work can't be cached ahead of time, because it depends on the query.\n\nAgents make this worse. A coding or research agent doesn't issue one query per task; it issues dozens, and each one pays the reranking bill again. In my own coding-agent sessions, I'd estimate roughly 80% of the work is finding the right context: searching, reading and deciding what matters. Finding context is becoming the inner loop of AI work.\n\nSo we asked a simple question: **how much of the reranker's job can the retrieval representation do by itself**, using only what was computed and stored before the query arrived?\n\n02 — The representation spectrum\n\n## One vector is too few. One per token is too many.\n\nThere have been two classic answers to the question *how should a document be stored for search?*\n\n**One vector per document.** Bi-encoders such as Qwen3-Embedding, E5 or EmbeddingGemma pool the whole text into a single point. It's cheap to store and fast to search. But every query is compared against the same summary: a question about the date and a question about the architect hit the same point in space. The score is one dot product, so there is nothing left to refine. That's why single vectors so often get paired with a reranker.\n\n**One vector per token.** Late-interaction models like ColBERT keep a contextual vector for every token and let each query token find its best match (MaxSim). That preserves detail and ranks well. The costs come with it: the index grows with every token stored, scoring cost grows with both query and document length, and all the scorer can do with a stored token vector is measure its similarity to a query token. Each query token takes its single best match; there is no separate content to *read*.\n\n**BreadBowl-Embed sits in between on purpose.** Every passage gets a fixed budget of 16 slots, however long it is. Each slot carries two vectors with two different jobs:\n\n- a routing vector (256-d) that says *where to look* . It is what gets indexed and searched.\n- a value vector (256-d) that holds *what to read* once a query has decided where to look.\n\n### What do you actually need from a representation?\n\nHere's Table 1 from our paper, made interactive. Switch on the requirements one at a time and watch which designs survive.\n\n| Architecture | Retriever | Reranker | Extra backbone pass | Separate value readout | Stored doc. vectors | Scoring cost | \n|---|---|---|---|---|---|---|\n| Bi-encoder | ✓ | ✓ <sup>*</sup> | None | — | 1 | *O* (*h* ) | \n| Cross-encoder | — | ✓ | Per pair | — | — | *B* (*n<sub>q</sub>* +*n<sub>d</sub>* ) | \n| Poly-encoder | — | ✓ | None | — | 1 | *O* (*m<sub>q</sub>**h* ) | \n| ColBERT | ✓ | ✓ | None | — | *n<sub>d</sub>* | *O* (*n<sub>q</sub>**n<sub>d</sub>**h* ) | \n| ConstBERT | ✓ | ✓ | None | — | *m<sub>d</sub>* | *O* (*n<sub>q</sub>**m<sub>d</sub>**h* ) | \n| MVA (pooled) | — | ✓ | None | ✓ | 2 *m<sub>d</sub>* | *O* (*m<sub>q</sub>**m<sub>d</sub>**h* ) | \n| BreadBowl-Embed | ✓ | ✓ | None | ✓ | 2 *m<sub>d</sub>* | *O* (*m<sub>q</sub>**m<sub>d</sub>**h* ) | \n\nAll 7 designs shown. Pick a requirement.\n\n*n*,\n\n<sub>q</sub>\n*n*: query and document token counts.\n\n<sub>d</sub>\n*m*,\n\n<sub>q</sub>\n*m*: pooled vectors per input (16 for BreadBowl-Embed).\n\n<sub>d</sub>\n*h*: vector width.\n\n*B*(ℓ): one backbone pass over ℓ tokens.\n\n<sup>*</sup>A bi-encoder \"reranks\" by reusing the same dot product it retrieved with. Costs describe candidate scoring only. From Table 1 of the paper.\n\nSwitch on all four and exactly one row is left; every other design misses at least one. Bi-encoders and ColBERT can retrieve, but their stored vectors can only be *matched*, never *read*. ConstBERT fixes ColBERT's storage growth but keeps MaxSim. Multi-Vector Attention (MVA) introduced the key–value readout we build on, but its evaluated pipeline only reranks candidates that another system retrieved. Cross-encoders remain the gold standard for precision, but they can't retrieve and must re-read text.\n\nBreadBowl-Embed's central idea is small to state: **make the vectors that find a document the same vectors that decide how to read it.** We call it *polymorphic routing*.\n\n03 — How it works · Encode once\n\n## Sixteen learned questions, asked of every passage.\n\nBreadBowl-Embed's Slot Encoder starts with an off-the-shelf language model: Qwen3.5-0.8B-Base reads the text and produces a contextual vector for every token. A ColBERT-style model would store those token vectors. We don't.\n\nInstead, it pools them with *Referenced Cross Attention* (RCA): a bank of 16 learned reference vectors, shared by every input, each attends over all the token states and pulls out one *slot*. You can think of the references as 16 learned questions the model asks of every passage. The questions are fixed; the answers depend on the text.\n\nTwo small heads then read each slot: a linear routing head, and a value head (linear, plus a small residual MLP for extra capacity). Both run once, at indexing time, and their outputs are stored side by side: 16 × 256 routing numbers and 16 × 256 value numbers per passage. That's 8,192 numbers whether the passage is 20 tokens long or 384.\n\n### What do the slots actually hold?\n\nBelow is the query–document pair we inspect in the paper's appendix. Pick any slot to see which tokens it pooled. Query slot 13 puts nearly half its weight on *completed*. Document slot 13 pools the date: the last digit of 1889, then *Fair*, then the 7 of 1887. Slot 14 grabs *Rome*. Nothing in training asks for slots like these; it's simply what this example looks like.\n\n04 — Retrieve with routing\n\n## Search with sixteen keys.\n\nAt query time, the query goes through the same encoder: one backbone pass over about twenty tokens. Each of its 16 routing vectors then searches an index holding every stored document routing vector. Hits are grouped by document into a shortlist, and the shortlist is rescored with the full routing score:\n\nIn words: for each query slot, take a soft maximum over the document's 16 slots, then average over query slots. As τ<sub>r</sub> → 0 this becomes mean MaxSim, the late-interaction score familiar from ColBERT, just over 16 slots instead of hundreds of tokens. In the paper's pipeline each query slot pulls its top 128 slot hits, those are grouped into a 1,024-document shortlist, and the best 100 by routing score go on to reranking.\n\n05 — Rerank by reading values\n\n## The same vectors get a second job.\n\nHere's the step a single-vector model can't take. For each candidate we already have its 16 × 16 routing similarities. Soften them with a higher temperature and they become attention weights over the candidate's *stored value vectors*:\n\nEach query slot gets its own readout: a weighted blend of the document's values, chosen by that query slot. The readout is compared with the query slot's own value vector, and the agreement is added to the routing score. There are no new projections at comparison time, no backbone and no text, just 16 × 16 attention over vectors that were already sitting in storage. The cost is independent of how long the candidate is.\n\nThe document never changes. *What the query reads from it does.*\n\nAsk about the Colosseum, and all 16 query slots put their heaviest weight on document slot 9, the one that pooled *Colosseum … is in Rome*. Ask which museum is in Paris, and all 16 move to slot 5, which pooled the Louvre sentence. The two Eiffel Tower questions mostly read the same tower slots (11 and 12), because they're about the same thing; the difference shows up in individual query slots. In the *completed* question, two of the query slots that pooled *completed* lean hardest on slot 13, the date slot. It's the same 8,192 stored numbers every time.\n\nA single-vector model can't do this: whatever you ask, it compares against the same point. ColBERT matches different tokens for different questions, but a match is all it can measure. BreadBowl-Embed adds a second channel: routing decides where to look, and values carry what's found there.\n\n06 — Does reading the values help?\n\n## Same candidates. Better order.\n\nThe cleanest test of the idea is to take one trained model and one candidate list, then rank the list two ways: routing alone (QK), or routing plus the value readout (QKV). Recall is identical by construction, so any change in ranking comes from reading the values.\n\nThe gains are not uniform. The largest come on question answering and fact-checking: NQ (+11.9), TREC-COVID (+11.9, on only 50 queries), FEVER (+8.7) and Climate-FEVER (+6.9). The largest losses come where relevance means *a similar or opposing text* rather than *contains the answer*: ArguAna's counter-argument retrieval (−5.2) and Quora's duplicate questions (−2.2). That fits the mechanism, but the pattern isn't clean: SciFact, also claim verification, slips (−1.05), and Touché-2020, argument retrieval like ArguAna, gains (+5.8). There's also a confound. NQ, FEVER, HotpotQA and MS MARCO have training splits in our distillation data, so BEIR isn't uniformly zero-shot for this model; TREC-COVID's +11.9 is the largest gain on a task with no training split there.\n\n### Where the top result changes\n\nA macro average can hide churn, so here's the blunt version. On our development set (5,000 queries searched over 3.41M documents), we counted the queries where reading the values moved a labeled passage *into* first place, and the ones where it pushed one *out*. This tally comes from the same runs as the paper's Tables 3 and 5, but it isn't in the paper.\n\nOverall, values put a labeled passage on top **479** times and knocked one off **238** times: a net gain of 241, with plenty of churn. FEVER is lopsided (178 fixes, 7 breaks). MS MARCO goes the other way (112 fixes, 126 breaks), PubMedQA is slightly negative (2 vs 4), and SQuAD v2 breaks even. By nDCG@10, values improve six of the nine sources (paper, Table 5).\n\n### How close is that to a real reranker?\n\nWe trained the value pathway by distilling Qwen3-Reranker-0.6B, a cross-encoder. On the fixed development candidate lists used for checkpoint selection, the teacher scores 69.80 nDCG@10, routing alone 60.36, and routing plus values 63.40. So on those lists the readout closes about a third of the gap to its teacher, without running a transformer pass at rerank time.\n\nIt doesn't replace a cross-encoder: two-thirds of that gap remains, and you can still run one on top. What it does is move a real share of the reranker's value into a step that costs about 0.27 MFLOPs per candidate.\n\n### Try it on real queries\n\nThese are real questions and fact-checking claims from an earlier validation set, each with a 10–12 passage shortlist taken from saved candidate lists and re-scored with the released weights. Flip between routing-only and routing-plus-values ordering and watch where the labeled passage lands. The last one is a case where values make things worse, because that happens too.\n\nLook at what routing alone puts near the top of the Natural Questions lists: title stubs and list fragments such as *Thursday Night Football* or *1. “Take It Easy”*. They match the topic, but there's nothing in them that answers the question, and their value scores come out negative, so they sink. The Hamilton claim shows another pattern: routing ranks several other Alexander Hamiltons first, and the values put the Founding Father's biography, the labeled evidence, on top. The miss is instructive too. For “The Rocket”, the values latch onto an answer-shaped passage (a roller coaster called The Rocket that was destroyed in 1979) about the wrong Rocket.\n\nRouting finds the topic. *Values look for the answer.*\n\n07 — Training\n\n## Teach it to rank. Replay how to find.\n\nSharing one representation between retrieval and reranking sets up a tug-of-war. Training the value pathway to imitate a cross-encoder updates the shared Slot Encoder, and that can quietly move the routing vectors that retrieval depends on. A sharper reranker over worse candidates isn't progress.\n\nDoes replay matter? We ran the same distillation with and without it: same initialization, same queries, 4,000 updates each, one seed per arm. Replay does add training compute (8.2M extra query–positive presentations). Each checkpoint then re-encoded the development corpora and retrieved its own candidates, so the comparison captures discovery as well as ordering.\n\n| Training | Recall@100 | nDCG@10 routing | nDCG@10 routing + values | \n|---|---|---|---|\n| KL distillation only | 86.63 | 58.64 | 62.57 | \n| KL + retrieval replay | 87.19 +0.56 | 60.40 +1.76 | 63.42 +0.85 | \n\nSource-macro averages over the 5,000-query development set (paper, Table 3). The replay gain in final nDCG@10 (0.84 before rounding) has a 95% bootstrap interval of [0.45, 1.24] points.\n\nReplay improves both what we find and how we order it. One honest wrinkle: recall of the coarse shortlist, before routing picks the final 100, dipped slightly (88.50 → 88.38), so replay doesn't preserve every aspect of candidate coverage.\n\n08 — Where it stands\n\n## Not the leaderboard leader. A new shape of model.\n\nWe want to be precise about what this release is and isn't. On published BEIR averages, BreadBowl-Embed trails strong single-vector models, including some much smaller ones:\n\nBecause published numbers aren't directly comparable, we also re-ran two of those models ourselves on identical development data, with the same queries, corpora, labels and token limits:\n\n| Model | nDCG@10 | Recall@100 | \n|---|---|---|\n| Qwen3-Embedding-0.6B | 62.25 | 87.42 | \n| Jina v5 text-small (retrieval) | 66.14 | 90.73 | \n| BreadBowl-Embed, routing only | 60.40 | 87.19 | \n| BreadBowl-Embed, routing + values | 63.42 | 87.19 | \n\nSource-macro averages over nine development sources (paper, Table 9). Each system retrieves its own candidates. Training data, model size and training compute are not matched. These queries come from the same sources as BreadBowl-Embed's training data, and the set was used for model selection.\n\nRouting alone trails Qwen3-Embedding-0.6B. Reading the values puts BreadBowl-Embed ahead on average, at similar recall, but the lead is concentrated: it wins four of the nine sources, and FEVER and NQ carry the average. Jina v5 text-small is ahead of both.\n\nSo why release now? Because the interesting result isn't the leaderboard position; it's the mechanism. Same model, same candidates, +3.07 nDCG@10 on BEIR, from information that was already sitting in the index. (That comparison switches the value score off at inference in one trained model; it isn't a separately trained routing-only model.) A single-vector index has no stored content to read at query time, so it has no equivalent step to take. Whether the value gain grows or shrinks as routing improves is an open question: in our replay ablation, better routing came with a smaller gain from values (3.93 → 3.02 points). And we haven't scaled any part of the recipe yet.\n\n- End-to-end latency and throughput against a retriever + cross-encoder stack. The efficiency argument here is architectural arithmetic, not a benchmark.\n- Cross-encoder-level precision. On our development lists, the 0.6B teacher is still 6.4 nDCG@10 points ahead.\n- A head-to-head with ColBERT-family models under matched training.\n- Variance across training seeds, or a full audit of overlap between training data and BEIR. Some BEIR domains appear in our adaptation data, and earlier SciFact and NFCorpus runs informed development.\n- Compatibility of stored embeddings across model versions.\n\n09 — Try it\n\n## Encode once. Retrieve and rerank.\n\nThe weights are on Hugging Face under Apache-2.0 as `breadbowl-embed-v1.2-preview`,<sup>6</sup> and the inference toolkit is on GitHub (it isn't on PyPI yet). It runs on CPU by default, and on Apple Silicon with `device=\"mps\"`. Python 3.12 or newer.\n\n`python -m pip install \"git+https://github.com/BreadBowlAI/breadbowl-embed.git\"````\nfrom breadbowl_embed import BreadBowl\nmodel = BreadBowl.from_pretrained(\"dotproductx/breadbowl-embed-v1.2-preview\")\npassages = [\n    \"Octopuses have three hearts.\",\n    \"The Pacific is Earth's largest ocean.\",\n    \"Plants use sunlight during photosynthesis.\",\n]\nfor hit in model.rerank(\"How many hearts does an octopus have?\", passages):\n    print(hit.score, passages[hit.document_index])\n```\nFor repeated queries, encode documents once and reuse the stored slots. The index returns routing, value and combined scores for every hit:\n\n```\ndocuments = model.encode_documents(passages, batch_size=2)   # routing + value slots, stored once\nindex = model.index(documents, ids=[\"octopus\", \"ocean\", \"plants\"])\nqueries = model.encode_queries([\"How many hearts does an octopus have?\"])\nfor hit in index.search(queries, top_k=2, candidate_k=3)[0]:\n    print(hit.id, hit.routing_score, hit.value_score, hit.score)\nprint(documents.routing.shape)  # [3, 16, 256]\nprint(documents.value.shape)    # [3, 16, 256]\n```\nThe bundled index is an exact linear scan over every document's routing slots, meant for experiments. It doesn't implement the paper's per-slot shortlist pipeline, and it has no ANN backend. Raw vectors take 32 KiB per passage in float32. Keep document values unnormalized before the attention read; normalizing them changes the scorer.\n\n10 — Why we're building this\n\n## Representations that can be *read*, not just matched.\n\n    More and more AI products are retrieval products underneath. Coding agents search repositories, enterprise copilots search company knowledge, research agents search the web, and recommendation systems search through people, products and intent. Underneath all of them sits a representation layer that, today, mostly points at nearest neighbors and leaves the real judgment to a second model.\n\nBreadBowl's bet is that this layer should carry more structure: enough to rank, to adapt to a domain, and to be read differently by every question. BreadBowl-Embed is the first step: **one stored representation that both finds and reads.**\n\n### Want to see what it finds in your data?\n\nAlpha access to BreadBowl's hosted service is open for model evaluations, design partnerships and teams building retrieval-heavy products. We'd especially love to hear from you if a reranker is your bottleneck.\n\ncontact@breadbowl.ai\n### Citation\n\nNotes\n\n1. Cross-encoder estimate: forward FLOPs ≈ 2 × non-embedding parameters × tokens. For Qwen3-Reranker-0.6B (our distillation teacher, ≈0.44B non-embedding parameters) on a ≈420-token (query + 384-token passage) pair, that is ≈0.4 TFLOPs per candidate once attention is included, or ≈40 TFLOPs for 100 candidates; the reranker's prompt template adds tokens, so this is conservative. Shorter passages shrink it proportionally. BreadBowl-Embed's value read: 2 × (16·16·256 + 16·16·256 + 16·256) ≈ 0.27 MFLOPs per candidate, or ≈27 MFLOPs for 100. Its first stage is heavier than a single-vector search, though: an exact flat search compares 16 query vectors with 16 vectors per document (65,536 multiply-adds per document, 64× a 1,024-d single vector), and rescoring a 1,024-document shortlist with full routing scores adds ≈0.13 GFLOPs. None of this includes moving stored vectors (16–32 KiB per candidate). We have not measured end-to-end latency.\n2. Counts follow Table 1 of the paper and precede compression and index duplication. ColBERT is shown with 128-d token vectors and 32-token queries; BreadBowl-Embed with 16 slots of 256 routing + 256 value dimensions. Scoring cost: ColBERT 32 × *n* × 128 multiply-adds; BreadBowl-Embed 16 × 16 × 256 (routing) + 16 × 16 × 256 (value read) + 16 × 256 (comparison) per passage. At 384 tokens a 2-bit ColBERTv2 index needs ≈13.5 KiB per passage, versus 16 KiB for BreadBowl-Embed's slots in bf16.\n3. Demo numbers come from running the released weights through a NumPy re-implementation of the inference path, in float32 on CPU. Its tokenization matches the reference token IDs, and its routing vectors agree with the reference GPU outputs (bf16) at a per-slot cosine of 0.9998 or better, so values can differ from the paper's figures in the third decimal.\n4. Official BEIR full corpora and relevance judgments, test splits except MS MARCO (dev). nDCG computed with pytrec_eval. CQADupStack's 12 forums are averaged into one task before the 15-task macro average. Candidates come from Algorithm 1 of the paper (128 slot hits per query slot, a 1,024-document shortlist, the top 100 by routing score) with exact flat search. One fixed checkpoint and one set of retrieval settings for every task; results are point estimates.\n5. 5,000 development queries from nine sources (LightOn subsets of MS MARCO, NQ, TriviaQA, SQuAD v2, HotpotQA and FEVER, plus StackExchange duplicates, SciRepEval search and PubMedQA), each searched over its own prepared corpus (3.41M documents in total). They were held out before candidate mining but come from the same sources as the training data, and they were used for model selection. The top-1 fix/break counts are our own tally from the same runs.\n6. The checkpoint is published as `breadbowl-embed-v1.2-preview` . \"v1.2\" is the internal research iteration; this is BreadBowl's first public model release.","body_html":"<p><strong>Introducing BreadBowl-Embed</strong>  ·  Open weights  ·  Apache-2.0</p>\n<h1 id=\"one-vector-is-too-few-one-per-token-is-too-many\">One vector is too <em>few</em>. One per token is too <em>many</em>.</h1>\n<pre><code>BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text.</code></pre>\n<p>01 — The problem</p>\n<h2 id=\"high-precision-retrieval-reads-your-documents-twice\">High-precision retrieval reads your documents twice.</h2>\n<p>Once to find them, and again to decide which ones actually matter.</p>\n<p>Modern search and RAG pipelines run in two stages. First, a <strong>bi-encoder</strong> compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance (the date, the exception, who did what to whom), a <strong>cross-encoder</strong> re-reads the top candidates next to the query and scores them again.</p>\n<p>That second stage is where the precision comes from, and it is expensive in a very specific way: the reranker runs a full transformer pass over every (query, candidate) pair, every time a query arrives. Most of that work can&#39;t be cached ahead of time, because it depends on the query.</p>\n<p>Agents make this worse. A coding or research agent doesn&#39;t issue one query per task; it issues dozens, and each one pays the reranking bill again. In my own coding-agent sessions, I&#39;d estimate roughly 80% of the work is finding the right context: searching, reading and deciding what matters. Finding context is becoming the inner loop of AI work.</p>\n<p>So we asked a simple question: <strong>how much of the reranker&#39;s job can the retrieval representation do by itself</strong>, using only what was computed and stored before the query arrived?</p>\n<p>02 — The representation spectrum</p>\n<h2 id=\"one-vector-is-too-few-one-per-token-is-too-many-2\">One vector is too few. One per token is too many.</h2>\n<p>There have been two classic answers to the question <em>how should a document be stored for search?</em></p>\n<p><strong>One vector per document.</strong> Bi-encoders such as Qwen3-Embedding, E5 or EmbeddingGemma pool the whole text into a single point. It&#39;s cheap to store and fast to search. But every query is compared against the same summary: a question about the date and a question about the architect hit the same point in space. The score is one dot product, so there is nothing left to refine. That&#39;s why single vectors so often get paired with a reranker.</p>\n<p><strong>One vector per token.</strong> Late-interaction models like ColBERT keep a contextual vector for every token and let each query token find its best match (MaxSim). That preserves detail and ranks well. The costs come with it: the index grows with every token stored, scoring cost grows with both query and document length, and all the scorer can do with a stored token vector is measure its similarity to a query token. Each query token takes its single best match; there is no separate content to <em>read</em>.</p>\n<p><strong>BreadBowl-Embed sits in between on purpose.</strong> Every passage gets a fixed budget of 16 slots, however long it is. Each slot carries two vectors with two different jobs:</p>\n<ul><li>a routing vector (256-d) that says <em>where to look</em> . It is what gets indexed and searched.</li><li>a value vector (256-d) that holds <em>what to read</em> once a query has decided where to look.</li></ul>\n<h3 id=\"what-do-you-actually-need-from-a-representation\">What do you actually need from a representation?</h3>\n<p>Here&#39;s Table 1 from our paper, made interactive. Switch on the requirements one at a time and watch which designs survive.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Architecture</th><th>Retriever</th><th>Reranker</th><th>Extra backbone pass</th><th>Separate value readout</th><th>Stored doc. vectors</th><th>Scoring cost</th></tr></thead><tbody><tr><td>Bi-encoder</td><td>✓</td><td>✓ &lt;sup&gt;*&lt;/sup&gt;</td><td>None</td><td>—</td><td>1</td><td>*O* (*h* )</td></tr><tr><td>Cross-encoder</td><td>—</td><td>✓</td><td>Per pair</td><td>—</td><td>—</td><td>*B* (<em>n&lt;sub&gt;q&lt;/sub&gt;</em> +<em>n&lt;sub&gt;d&lt;/sub&gt;</em> )</td></tr><tr><td>Poly-encoder</td><td>—</td><td>✓</td><td>None</td><td>—</td><td>1</td><td>*O* (<em>m&lt;sub&gt;q&lt;/sub&gt;</em>*h* )</td></tr><tr><td>ColBERT</td><td>✓</td><td>✓</td><td>None</td><td>—</td><td><em>n&lt;sub&gt;d&lt;/sub&gt;</em></td><td>*O* (<em>n&lt;sub&gt;q&lt;/sub&gt;<strong>n&lt;sub&gt;d&lt;/sub&gt;</strong>h</em> )</td></tr><tr><td>ConstBERT</td><td>✓</td><td>✓</td><td>None</td><td>—</td><td><em>m&lt;sub&gt;d&lt;/sub&gt;</em></td><td>*O* (<em>n&lt;sub&gt;q&lt;/sub&gt;<strong>m&lt;sub&gt;d&lt;/sub&gt;</strong>h</em> )</td></tr><tr><td>MVA (pooled)</td><td>—</td><td>✓</td><td>None</td><td>✓</td><td>2 <em>m&lt;sub&gt;d&lt;/sub&gt;</em></td><td>*O* (<em>m&lt;sub&gt;q&lt;/sub&gt;<strong>m&lt;sub&gt;d&lt;/sub&gt;</strong>h</em> )</td></tr><tr><td>BreadBowl-Embed</td><td>✓</td><td>✓</td><td>None</td><td>✓</td><td>2 <em>m&lt;sub&gt;d&lt;/sub&gt;</em></td><td>*O* (<em>m&lt;sub&gt;q&lt;/sub&gt;<strong>m&lt;sub&gt;d&lt;/sub&gt;</strong>h</em> )</td></tr></tbody></table></div>\n<p>All 7 designs shown. Pick a requirement.</p>\n<p>*n*,</p>\n<p>&lt;sub&gt;q&lt;/sub&gt;\n*n*: query and document token counts.</p>\n<p>&lt;sub&gt;d&lt;/sub&gt;\n*m*,</p>\n<p>&lt;sub&gt;q&lt;/sub&gt;\n*m*: pooled vectors per input (16 for BreadBowl-Embed).</p>\n<p>&lt;sub&gt;d&lt;/sub&gt;\n*h*: vector width.</p>\n<p>*B*(ℓ): one backbone pass over ℓ tokens.</p>\n<p>&lt;sup&gt;*&lt;/sup&gt;A bi-encoder &quot;reranks&quot; by reusing the same dot product it retrieved with. Costs describe candidate scoring only. From Table 1 of the paper.</p>\n<p>Switch on all four and exactly one row is left; every other design misses at least one. Bi-encoders and ColBERT can retrieve, but their stored vectors can only be <em>matched</em>, never <em>read</em>. ConstBERT fixes ColBERT&#39;s storage growth but keeps MaxSim. Multi-Vector Attention (MVA) introduced the key–value readout we build on, but its evaluated pipeline only reranks candidates that another system retrieved. Cross-encoders remain the gold standard for precision, but they can&#39;t retrieve and must re-read text.</p>\n<p>BreadBowl-Embed&#39;s central idea is small to state: <strong>make the vectors that find a document the same vectors that decide how to read it.</strong> We call it <em>polymorphic routing</em>.</p>\n<p>03 — How it works · Encode once</p>\n<h2 id=\"sixteen-learned-questions-asked-of-every-passage\">Sixteen learned questions, asked of every passage.</h2>\n<p>BreadBowl-Embed&#39;s Slot Encoder starts with an off-the-shelf language model: Qwen3.5-0.8B-Base reads the text and produces a contextual vector for every token. A ColBERT-style model would store those token vectors. We don&#39;t.</p>\n<p>Instead, it pools them with <em>Referenced Cross Attention</em> (RCA): a bank of 16 learned reference vectors, shared by every input, each attends over all the token states and pulls out one <em>slot</em>. You can think of the references as 16 learned questions the model asks of every passage. The questions are fixed; the answers depend on the text.</p>\n<p>Two small heads then read each slot: a linear routing head, and a value head (linear, plus a small residual MLP for extra capacity). Both run once, at indexing time, and their outputs are stored side by side: 16 × 256 routing numbers and 16 × 256 value numbers per passage. That&#39;s 8,192 numbers whether the passage is 20 tokens long or 384.</p>\n<h3 id=\"what-do-the-slots-actually-hold\">What do the slots actually hold?</h3>\n<p>Below is the query–document pair we inspect in the paper&#39;s appendix. Pick any slot to see which tokens it pooled. Query slot 13 puts nearly half its weight on <em>completed</em>. Document slot 13 pools the date: the last digit of 1889, then <em>Fair</em>, then the 7 of 1887. Slot 14 grabs <em>Rome</em>. Nothing in training asks for slots like these; it&#39;s simply what this example looks like.</p>\n<p>04 — Retrieve with routing</p>\n<h2 id=\"search-with-sixteen-keys\">Search with sixteen keys.</h2>\n<p>At query time, the query goes through the same encoder: one backbone pass over about twenty tokens. Each of its 16 routing vectors then searches an index holding every stored document routing vector. Hits are grouped by document into a shortlist, and the shortlist is rescored with the full routing score:</p>\n<p>In words: for each query slot, take a soft maximum over the document&#39;s 16 slots, then average over query slots. As τ&lt;sub&gt;r&lt;/sub&gt; → 0 this becomes mean MaxSim, the late-interaction score familiar from ColBERT, just over 16 slots instead of hundreds of tokens. In the paper&#39;s pipeline each query slot pulls its top 128 slot hits, those are grouped into a 1,024-document shortlist, and the best 100 by routing score go on to reranking.</p>\n<p>05 — Rerank by reading values</p>\n<h2 id=\"the-same-vectors-get-a-second-job\">The same vectors get a second job.</h2>\n<p>Here&#39;s the step a single-vector model can&#39;t take. For each candidate we already have its 16 × 16 routing similarities. Soften them with a higher temperature and they become attention weights over the candidate&#39;s <em>stored value vectors</em>:</p>\n<p>Each query slot gets its own readout: a weighted blend of the document&#39;s values, chosen by that query slot. The readout is compared with the query slot&#39;s own value vector, and the agreement is added to the routing score. There are no new projections at comparison time, no backbone and no text, just 16 × 16 attention over vectors that were already sitting in storage. The cost is independent of how long the candidate is.</p>\n<p>The document never changes. <em>What the query reads from it does.</em></p>\n<p>Ask about the Colosseum, and all 16 query slots put their heaviest weight on document slot 9, the one that pooled <em>Colosseum … is in Rome</em>. Ask which museum is in Paris, and all 16 move to slot 5, which pooled the Louvre sentence. The two Eiffel Tower questions mostly read the same tower slots (11 and 12), because they&#39;re about the same thing; the difference shows up in individual query slots. In the <em>completed</em> question, two of the query slots that pooled <em>completed</em> lean hardest on slot 13, the date slot. It&#39;s the same 8,192 stored numbers every time.</p>\n<p>A single-vector model can&#39;t do this: whatever you ask, it compares against the same point. ColBERT matches different tokens for different questions, but a match is all it can measure. BreadBowl-Embed adds a second channel: routing decides where to look, and values carry what&#39;s found there.</p>\n<p>06 — Does reading the values help?</p>\n<h2 id=\"same-candidates-better-order\">Same candidates. Better order.</h2>\n<p>The cleanest test of the idea is to take one trained model and one candidate list, then rank the list two ways: routing alone (QK), or routing plus the value readout (QKV). Recall is identical by construction, so any change in ranking comes from reading the values.</p>\n<p>The gains are not uniform. The largest come on question answering and fact-checking: NQ (+11.9), TREC-COVID (+11.9, on only 50 queries), FEVER (+8.7) and Climate-FEVER (+6.9). The largest losses come where relevance means <em>a similar or opposing text</em> rather than <em>contains the answer</em>: ArguAna&#39;s counter-argument retrieval (−5.2) and Quora&#39;s duplicate questions (−2.2). That fits the mechanism, but the pattern isn&#39;t clean: SciFact, also claim verification, slips (−1.05), and Touché-2020, argument retrieval like ArguAna, gains (+5.8). There&#39;s also a confound. NQ, FEVER, HotpotQA and MS MARCO have training splits in our distillation data, so BEIR isn&#39;t uniformly zero-shot for this model; TREC-COVID&#39;s +11.9 is the largest gain on a task with no training split there.</p>\n<h3 id=\"where-the-top-result-changes\">Where the top result changes</h3>\n<p>A macro average can hide churn, so here&#39;s the blunt version. On our development set (5,000 queries searched over 3.41M documents), we counted the queries where reading the values moved a labeled passage <em>into</em> first place, and the ones where it pushed one <em>out</em>. This tally comes from the same runs as the paper&#39;s Tables 3 and 5, but it isn&#39;t in the paper.</p>\n<p>Overall, values put a labeled passage on top <strong>479</strong> times and knocked one off <strong>238</strong> times: a net gain of 241, with plenty of churn. FEVER is lopsided (178 fixes, 7 breaks). MS MARCO goes the other way (112 fixes, 126 breaks), PubMedQA is slightly negative (2 vs 4), and SQuAD v2 breaks even. By nDCG@10, values improve six of the nine sources (paper, Table 5).</p>\n<h3 id=\"how-close-is-that-to-a-real-reranker\">How close is that to a real reranker?</h3>\n<p>We trained the value pathway by distilling Qwen3-Reranker-0.6B, a cross-encoder. On the fixed development candidate lists used for checkpoint selection, the teacher scores 69.80 nDCG@10, routing alone 60.36, and routing plus values 63.40. So on those lists the readout closes about a third of the gap to its teacher, without running a transformer pass at rerank time.</p>\n<p>It doesn&#39;t replace a cross-encoder: two-thirds of that gap remains, and you can still run one on top. What it does is move a real share of the reranker&#39;s value into a step that costs about 0.27 MFLOPs per candidate.</p>\n<h3 id=\"try-it-on-real-queries\">Try it on real queries</h3>\n<p>These are real questions and fact-checking claims from an earlier validation set, each with a 10–12 passage shortlist taken from saved candidate lists and re-scored with the released weights. Flip between routing-only and routing-plus-values ordering and watch where the labeled passage lands. The last one is a case where values make things worse, because that happens too.</p>\n<p>Look at what routing alone puts near the top of the Natural Questions lists: title stubs and list fragments such as <em>Thursday Night Football</em> or <em>1. “Take It Easy”</em>. They match the topic, but there&#39;s nothing in them that answers the question, and their value scores come out negative, so they sink. The Hamilton claim shows another pattern: routing ranks several other Alexander Hamiltons first, and the values put the Founding Father&#39;s biography, the labeled evidence, on top. The miss is instructive too. For “The Rocket”, the values latch onto an answer-shaped passage (a roller coaster called The Rocket that was destroyed in 1979) about the wrong Rocket.</p>\n<p>Routing finds the topic. <em>Values look for the answer.</em></p>\n<p>07 — Training</p>\n<h2 id=\"teach-it-to-rank-replay-how-to-find\">Teach it to rank. Replay how to find.</h2>\n<p>Sharing one representation between retrieval and reranking sets up a tug-of-war. Training the value pathway to imitate a cross-encoder updates the shared Slot Encoder, and that can quietly move the routing vectors that retrieval depends on. A sharper reranker over worse candidates isn&#39;t progress.</p>\n<p>Does replay matter? We ran the same distillation with and without it: same initialization, same queries, 4,000 updates each, one seed per arm. Replay does add training compute (8.2M extra query–positive presentations). Each checkpoint then re-encoded the development corpora and retrieved its own candidates, so the comparison captures discovery as well as ordering.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Training</th><th>Recall@100</th><th>nDCG@10 routing</th><th>nDCG@10 routing + values</th></tr></thead><tbody><tr><td>KL distillation only</td><td>86.63</td><td>58.64</td><td>62.57</td></tr><tr><td>KL + retrieval replay</td><td>87.19 +0.56</td><td>60.40 +1.76</td><td>63.42 +0.85</td></tr></tbody></table></div>\n<p>Source-macro averages over the 5,000-query development set (paper, Table 3). The replay gain in final nDCG@10 (0.84 before rounding) has a 95% bootstrap interval of [0.45, 1.24] points.</p>\n<p>Replay improves both what we find and how we order it. One honest wrinkle: recall of the coarse shortlist, before routing picks the final 100, dipped slightly (88.50 → 88.38), so replay doesn&#39;t preserve every aspect of candidate coverage.</p>\n<p>08 — Where it stands</p>\n<h2 id=\"not-the-leaderboard-leader-a-new-shape-of-model\">Not the leaderboard leader. A new shape of model.</h2>\n<p>We want to be precise about what this release is and isn&#39;t. On published BEIR averages, BreadBowl-Embed trails strong single-vector models, including some much smaller ones:</p>\n<p>Because published numbers aren&#39;t directly comparable, we also re-ran two of those models ourselves on identical development data, with the same queries, corpora, labels and token limits:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>nDCG@10</th><th>Recall@100</th></tr></thead><tbody><tr><td>Qwen3-Embedding-0.6B</td><td>62.25</td><td>87.42</td></tr><tr><td>Jina v5 text-small (retrieval)</td><td>66.14</td><td>90.73</td></tr><tr><td>BreadBowl-Embed, routing only</td><td>60.40</td><td>87.19</td></tr><tr><td>BreadBowl-Embed, routing + values</td><td>63.42</td><td>87.19</td></tr></tbody></table></div>\n<p>Source-macro averages over nine development sources (paper, Table 9). Each system retrieves its own candidates. Training data, model size and training compute are not matched. These queries come from the same sources as BreadBowl-Embed&#39;s training data, and the set was used for model selection.</p>\n<p>Routing alone trails Qwen3-Embedding-0.6B. Reading the values puts BreadBowl-Embed ahead on average, at similar recall, but the lead is concentrated: it wins four of the nine sources, and FEVER and NQ carry the average. Jina v5 text-small is ahead of both.</p>\n<p>So why release now? Because the interesting result isn&#39;t the leaderboard position; it&#39;s the mechanism. Same model, same candidates, +3.07 nDCG@10 on BEIR, from information that was already sitting in the index. (That comparison switches the value score off at inference in one trained model; it isn&#39;t a separately trained routing-only model.) A single-vector index has no stored content to read at query time, so it has no equivalent step to take. Whether the value gain grows or shrinks as routing improves is an open question: in our replay ablation, better routing came with a smaller gain from values (3.93 → 3.02 points). And we haven&#39;t scaled any part of the recipe yet.</p>\n<ul><li>End-to-end latency and throughput against a retriever + cross-encoder stack. The efficiency argument here is architectural arithmetic, not a benchmark.</li><li>Cross-encoder-level precision. On our development lists, the 0.6B teacher is still 6.4 nDCG@10 points ahead.</li><li>A head-to-head with ColBERT-family models under matched training.</li><li>Variance across training seeds, or a full audit of overlap between training data and BEIR. Some BEIR domains appear in our adaptation data, and earlier SciFact and NFCorpus runs informed development.</li><li>Compatibility of stored embeddings across model versions.</li></ul>\n<p>09 — Try it</p>\n<h2 id=\"encode-once-retrieve-and-rerank\">Encode once. Retrieve and rerank.</h2>\n<p>The weights are on Hugging Face under Apache-2.0 as <code>breadbowl-embed-v1.2-preview</code>,&lt;sup&gt;6&lt;/sup&gt; and the inference toolkit is on GitHub (it isn&#39;t on PyPI yet). It runs on CPU by default, and on Apple Silicon with <code>device=&quot;mps&quot;</code>. Python 3.12 or newer.</p>\n<p><code>python -m pip install &quot;git+https://github.com/BreadBowlAI/breadbowl-embed.git&quot;</code>```\nfrom breadbowl_embed import BreadBowl\nmodel = BreadBowl.from_pretrained(&quot;dotproductx/breadbowl-embed-v1.2-preview&quot;)\npassages = [\n    &quot;Octopuses have three hearts.&quot;,\n    &quot;The Pacific is Earth&#39;s largest ocean.&quot;,\n    &quot;Plants use sunlight during photosynthesis.&quot;,\n]\nfor hit in model.rerank(&quot;How many hearts does an octopus have?&quot;, passages):\n    print(hit.score, passages[hit.document_index])</p>\n<pre><code>For repeated queries, encode documents once and reuse the stored slots. The index returns routing, value and combined scores for every hit:\n</code></pre>\n<p>documents = model.encode_documents(passages, batch_size=2)   # routing + value slots, stored once\nindex = model.index(documents, ids=[&quot;octopus&quot;, &quot;ocean&quot;, &quot;plants&quot;])\nqueries = model.encode_queries([&quot;How many hearts does an octopus have?&quot;])\nfor hit in index.search(queries, top_k=2, candidate_k=3)[0]:\n    print(hit.id, hit.routing_score, hit.value_score, hit.score)\nprint(documents.routing.shape)  # [3, 16, 256]\nprint(documents.value.shape)    # [3, 16, 256]</p>\n<pre><code>The bundled index is an exact linear scan over every document&#39;s routing slots, meant for experiments. It doesn&#39;t implement the paper&#39;s per-slot shortlist pipeline, and it has no ANN backend. Raw vectors take 32 KiB per passage in float32. Keep document values unnormalized before the attention read; normalizing them changes the scorer.\n\n10 — Why we&#39;re building this\n\n## Representations that can be *read*, not just matched.\n\n    More and more AI products are retrieval products underneath. Coding agents search repositories, enterprise copilots search company knowledge, research agents search the web, and recommendation systems search through people, products and intent. Underneath all of them sits a representation layer that, today, mostly points at nearest neighbors and leaves the real judgment to a second model.\n\nBreadBowl&#39;s bet is that this layer should carry more structure: enough to rank, to adapt to a domain, and to be read differently by every question. BreadBowl-Embed is the first step: **one stored representation that both finds and reads.**\n\n### Want to see what it finds in your data?\n\nAlpha access to BreadBowl&#39;s hosted service is open for model evaluations, design partnerships and teams building retrieval-heavy products. We&#39;d especially love to hear from you if a reranker is your bottleneck.\n\ncontact@breadbowl.ai\n### Citation\n\nNotes\n\n1. Cross-encoder estimate: forward FLOPs ≈ 2 × non-embedding parameters × tokens. For Qwen3-Reranker-0.6B (our distillation teacher, ≈0.44B non-embedding parameters) on a ≈420-token (query + 384-token passage) pair, that is ≈0.4 TFLOPs per candidate once attention is included, or ≈40 TFLOPs for 100 candidates; the reranker&#39;s prompt template adds tokens, so this is conservative. Shorter passages shrink it proportionally. BreadBowl-Embed&#39;s value read: 2 × (16·16·256 + 16·16·256 + 16·256) ≈ 0.27 MFLOPs per candidate, or ≈27 MFLOPs for 100. Its first stage is heavier than a single-vector search, though: an exact flat search compares 16 query vectors with 16 vectors per document (65,536 multiply-adds per document, 64× a 1,024-d single vector), and rescoring a 1,024-document shortlist with full routing scores adds ≈0.13 GFLOPs. None of this includes moving stored vectors (16–32 KiB per candidate). We have not measured end-to-end latency.\n2. Counts follow Table 1 of the paper and precede compression and index duplication. ColBERT is shown with 128-d token vectors and 32-token queries; BreadBowl-Embed with 16 slots of 256 routing + 256 value dimensions. Scoring cost: ColBERT 32 × *n* × 128 multiply-adds; BreadBowl-Embed 16 × 16 × 256 (routing) + 16 × 16 × 256 (value read) + 16 × 256 (comparison) per passage. At 384 tokens a 2-bit ColBERTv2 index needs ≈13.5 KiB per passage, versus 16 KiB for BreadBowl-Embed&#39;s slots in bf16.\n3. Demo numbers come from running the released weights through a NumPy re-implementation of the inference path, in float32 on CPU. Its tokenization matches the reference token IDs, and its routing vectors agree with the reference GPU outputs (bf16) at a per-slot cosine of 0.9998 or better, so values can differ from the paper&#39;s figures in the third decimal.\n4. Official BEIR full corpora and relevance judgments, test splits except MS MARCO (dev). nDCG computed with pytrec_eval. CQADupStack&#39;s 12 forums are averaged into one task before the 15-task macro average. Candidates come from Algorithm 1 of the paper (128 slot hits per query slot, a 1,024-document shortlist, the top 100 by routing score) with exact flat search. One fixed checkpoint and one set of retrieval settings for every task; results are point estimates.\n5. 5,000 development queries from nine sources (LightOn subsets of MS MARCO, NQ, TriviaQA, SQuAD v2, HotpotQA and FEVER, plus StackExchange duplicates, SciRepEval search and PubMedQA), each searched over its own prepared corpus (3.41M documents in total). They were held out before candidate mining but come from the same sources as the training data, and they were used for model selection. The top-1 fix/break counts are our own tally from the same runs.\n6. The checkpoint is published as `breadbowl-embed-v1.2-preview` . &quot;v1.2&quot; is the internal research iteration; this is BreadBowl&#39;s first public model release.</code></pre>","headings":[{"level":1,"text":"One vector is too *few*. One per token is too *many*.","id":"one-vector-is-too-few-one-per-token-is-too-many"},{"level":2,"text":"High-precision retrieval reads your documents twice.","id":"high-precision-retrieval-reads-your-documents-twice"},{"level":2,"text":"One vector is too few. One per token is too many.","id":"one-vector-is-too-few-one-per-token-is-too-many-2"},{"level":3,"text":"What do you actually need from a representation?","id":"what-do-you-actually-need-from-a-representation"},{"level":2,"text":"Sixteen learned questions, asked of every passage.","id":"sixteen-learned-questions-asked-of-every-passage"},{"level":3,"text":"What do the slots actually hold?","id":"what-do-the-slots-actually-hold"},{"level":2,"text":"Search with sixteen keys.","id":"search-with-sixteen-keys"},{"level":2,"text":"The same vectors get a second job.","id":"the-same-vectors-get-a-second-job"},{"level":2,"text":"Same candidates. Better order.","id":"same-candidates-better-order"},{"level":3,"text":"Where the top result changes","id":"where-the-top-result-changes"},{"level":3,"text":"How close is that to a real reranker?","id":"how-close-is-that-to-a-real-reranker"},{"level":3,"text":"Try it on real queries","id":"try-it-on-real-queries"},{"level":2,"text":"Teach it to rank. Replay how to find.","id":"teach-it-to-rank-replay-how-to-find"},{"level":2,"text":"Not the leaderboard leader. A new shape of model.","id":"not-the-leaderboard-leader-a-new-shape-of-model"},{"level":2,"text":"Encode once. Retrieve and rerank.","id":"encode-once-retrieve-and-rerank"}]}}