{"article":{"slug":"rip-vector-database","title":"RIP, vector database","subtitle":null,"summary":"turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.","content_type":"blog_post","language":"en","canonical_url":"https://turbopuffer.com/blog/rip-vector-database","author":{"name":null,"url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"turbopuffer","url":"https://turbopuffer.com/","listing_slug":null,"listing":null},"topics":[{"name":"Databases","slug":"databases","url":"https://listedarticles.com/topics/databases"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":"https://turbopuffer.com/og/turbopuffer.png","license":"all-rights-reserved","word_count":1334,"reading_minutes":6,"published_at":"2026-10-01T18:09:54.403Z","added_at":"2026-10-01T18:09:54.403Z","updated_at":"2026-10-01T18:09:54.403Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/rip-vector-database","markdown_url":"https://listedarticles.com/articles/rip-vector-database.md","example":false,"citation":"turbopuffer. \"RIP, vector database.\" 1 Oct 2026. https://turbopuffer.com/blog/rip-vector-database (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://turbopuffer.com/blog/rip-vector-database"},"body_markdown":"We are changing turbopuffer's storage architecture to take search to the next level. turbopuffer v3 changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to make search faster in every respect, including vector search, but it also lays the foundation to move many more SQL queries to turbopuffer and make them fast.\n\nturbopuffer launched as a serverless vector database (v1), highly specialized to the task of serving extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion.\n\nturbopuffer evolved to have very strong text and regex search (v2), and is being\nused for many non-search use cases, like\nLinear's syncing engine.\nThe query engine has evolved along the way to support all of these query plans,\nbut the storage architecture has remained largely unchanged: the ANN vector\nindex was and still is the primary index around which all other indexes and\nquery plans revolve. This design has constrained several query plans, like\n`GROUP BY` and aggregations.\n\nWe've pushed the vector-primary architecture as far as we can, and it's time to move on. We're in the process of moving to a new primary index, and making ANN \"just another\" secondary index. We thought it might be fun to open up the doors and let you follow along.\n\nFor this first update, we'll set the stage with why we're doing this in the first place. Walk with me on a short journey from tpuf v1 to today.\n\nIn the first version of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing wisdom at the time was graph-based vector indexes, but a hierarchical clustering index plays better with object storage. We started with SPANN, and eventually migrated to SPFresh to support incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree with a single root.\n\n```\n      ┌───────────────┐\n      │ root centroid │\n      └───────────────┘\n        ╱     │     ╲\n       ╱      │      ╲\n┌────────┐┌────────┐┌────────┐\n│  leaf  ││  leaf  ││  leaf  │\n│centroid││centroid││centroid│\n└────────┘└────────┘└────────┘\n   ╱  ╲      ╱  ╲      ╱  ╲\n┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐\n│vec││vec││vec││vec││vec││vec│\n└───┘└───┘└───┘└───┘└───┘└───┘\n```\nWe implemented this on top of a storage layer presenting as a key-value map,\nwith sorted and unique keys. Each cluster is given a `ClusterId`, and vectors\nwithin each cluster are given a dense `LocalId`.\n\n```\n// leaf vectors\nK::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Vector(C0L1) = vec![-0.28, 0.96, ...]\nK::Id(C0L1) = 13\n// cluster centroid for C0 is itself clustered at the next level of the tree\nK::Vector(C1L4) = vec![0.64, -0.48, ...]\nK::Id(C1L4) = C0\n```\nAs you can see above, everything is keyed by `ClusterId` and `LocalId` (e.g.\n`C0L1`), which together we call the **ANN address**. This is what we mean when\nwe say the ANN index is the primary index.\n\nTwo new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search.\n\nNaturally, customers wanted to be able to add attribute values and filter vector searches on them. To make filtering fast and high-recall, we modeled these as an inverted index that maps an attribute value to the ANN address of the documents that contain it.\n\n```\nK::AttrIndex(\"family\", \"Alcidae\") -> vec![C0L3, C1L2, C1L3, ...]\nK::AttrIndex(\"genus\", \"Fratercula\") -> vec![C0L3, C1L2, C1L9, ...]\n```\nFor projections (`include_attributes`), we also stored the document attributes\nalongside the ID and the vector.\n\n```\nK::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Attr(C0L0, \"family\") = \"Alcidae\"\nK::Attr(C0L0, \"genus\") = \"Fratercula\"\n```\nBM25 full-text search was another obvious and much-demanded query plan. Similar\nto attribute search, full-text search works by first finding the documents that\nhave the query term present (commonly called \"postings\"). For an FTS index, we\nalso include the `(term count, document length)` metadata necessary for BM25\nscoring:\n\n```\nK::FTS(\"description\", \"Atlantic\") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]\nK::Attr(C0L0, \"description\") -> \"A sharply dressed black-and-white seabird with a \\\nhuge, multicolored bill, the Atlantic Puffin is often \\\ncalled the clown of the sea. It breeds in burrows on \\\nislands in the North Atlantic, and winters at sea.\"\n```\nOver time, we've shipped several other index structures and query engines: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built around the same vector-primary storage layout.\n\nThe ANN primary index has largely remained intact until today for one simple reason: it works really, really well for ANN search on object storage. On top of this architecture, we've pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. Any significant change here risks introducing regressions in ANN performance.\n\nHowever, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization.\n\nAs described above, turbopuffer currently puts the full contents of each document under its ANN address. When there is only one vector, the non-vector data is stored alongside the vector only once.\n\nHowever, for multi-vector representations of a document, such as document nesting or late interaction, this means we have to duplicate the contents for each vector. This is the reason for some of our more unfortunate limits.\n\nAny time a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to ensure they remain well clustered (otherwise recall may suffer). Because everything in a document is stored keyed by the ANN address of the document's vector, this rebalancing cascades to moving the full document contents, as well as any inverted (attribute and FTS) indexes that reference it. Updating just one vector can move hundreds of attributes and their indexes.\n\nThis write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.\n\nModern query engines are vectorized: they run tight loops over blocks of values, which amortizes fixed per-block costs, compresses better, keeps the CPU pipeline full, and unlocks SIMD. DuckDB, for example, works in batches of 2,048 rows, ClickHouse up to ~65k, Lucene's posting blocks are 256 docs, and our ANN index works best with clusters of around 100–200 documents. Every query plan has an optimal block size, but today they are all constrained by the ANN primary index. A plan that wants blocks of thousands of documents to keep the CPU saturated is still stuck at 100–200.\n\nWe've already documented how much this matters in turbopuffer. Our first version of full-text search partitioned posting lists along ANN cluster boundaries, and the median block held just ~1.5 postings. FTS v2 reworked postings into fixed blocks of ~256, and the index got 10x smaller and queries got up to 20x faster. Posting lists could do that because they're stored separately and point at documents, so their layout doesn't have to follow the clusters. Aggregations and other scans read the documents themselves, and those are stored one block per cluster. As long as the ANN address is the primary key, their block size is constrained to the cluster size, even if they'd prefer something larger.\n\nThe solution to these problems is simple: don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.\n\nv3 is a new foundation that will unlock significant performance improvement on\nall query plans, and we hit a major milestone earlier this month: 100% of CI\npasses on turbopuffer v3. We started by focusing on correctness. Now we will\nmake it correct *and* fast. Watching benchmark numbers go down is great fun, so\nwe wanted to get you in at day zero of perf grinding. We will share\nthe benchmarks in public over the coming weeks, as we work toward (and\nbeyond) performance parity before rolling out v3 to production.\n\nturbopuffer is a fast search engine that hosts 1T+ documents, handles 10M+ writes/s, and serves 25k+ queries/s. We are ready for far more. We hope you'll trust us with your queries.\n\nGet started","body_html":"<p>We are changing turbopuffer&#39;s storage architecture to take search to the next level. turbopuffer v3 changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to make search faster in every respect, including vector search, but it also lays the foundation to move many more SQL queries to turbopuffer and make them fast.</p>\n<p>turbopuffer launched as a serverless vector database (v1), highly specialized to the task of serving extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion.</p>\n<p>turbopuffer evolved to have very strong text and regex search (v2), and is being\nused for many non-search use cases, like\nLinear&#39;s syncing engine.\nThe query engine has evolved along the way to support all of these query plans,\nbut the storage architecture has remained largely unchanged: the ANN vector\nindex was and still is the primary index around which all other indexes and\nquery plans revolve. This design has constrained several query plans, like\n<code>GROUP BY</code> and aggregations.</p>\n<p>We&#39;ve pushed the vector-primary architecture as far as we can, and it&#39;s time to move on. We&#39;re in the process of moving to a new primary index, and making ANN &quot;just another&quot; secondary index. We thought it might be fun to open up the doors and let you follow along.</p>\n<p>For this first update, we&#39;ll set the stage with why we&#39;re doing this in the first place. Walk with me on a short journey from tpuf v1 to today.</p>\n<p>In the first version of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing wisdom at the time was graph-based vector indexes, but a hierarchical clustering index plays better with object storage. We started with SPANN, and eventually migrated to SPFresh to support incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree with a single root.</p>\n<pre><code>      ┌───────────────┐\n      │ root centroid │\n      └───────────────┘\n        ╱     │     ╲\n       ╱      │      ╲\n┌────────┐┌────────┐┌────────┐\n│  leaf  ││  leaf  ││  leaf  │\n│centroid││centroid││centroid│\n└────────┘└────────┘└────────┘\n   ╱  ╲      ╱  ╲      ╱  ╲\n┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐\n│vec││vec││vec││vec││vec││vec│\n└───┘└───┘└───┘└───┘└───┘└───┘</code></pre>\n<p>We implemented this on top of a storage layer presenting as a key-value map,\nwith sorted and unique keys. Each cluster is given a <code>ClusterId</code>, and vectors\nwithin each cluster are given a dense <code>LocalId</code>.</p>\n<pre><code>// leaf vectors\nK::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Vector(C0L1) = vec![-0.28, 0.96, ...]\nK::Id(C0L1) = 13\n// cluster centroid for C0 is itself clustered at the next level of the tree\nK::Vector(C1L4) = vec![0.64, -0.48, ...]\nK::Id(C1L4) = C0</code></pre>\n<p>As you can see above, everything is keyed by <code>ClusterId</code> and <code>LocalId</code> (e.g.\n<code>C0L1</code>), which together we call the <strong>ANN address</strong>. This is what we mean when\nwe say the ANN index is the primary index.</p>\n<p>Two new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search.</p>\n<p>Naturally, customers wanted to be able to add attribute values and filter vector searches on them. To make filtering fast and high-recall, we modeled these as an inverted index that maps an attribute value to the ANN address of the documents that contain it.</p>\n<pre><code>K::AttrIndex(&quot;family&quot;, &quot;Alcidae&quot;) -&gt; vec![C0L3, C1L2, C1L3, ...]\nK::AttrIndex(&quot;genus&quot;, &quot;Fratercula&quot;) -&gt; vec![C0L3, C1L2, C1L9, ...]</code></pre>\n<p>For projections (<code>include_attributes</code>), we also stored the document attributes\nalongside the ID and the vector.</p>\n<pre><code>K::Vector(C0L0) = vec![0.45, 0.32, ...]\nK::Id(C0L0) = 7\nK::Attr(C0L0, &quot;family&quot;) = &quot;Alcidae&quot;\nK::Attr(C0L0, &quot;genus&quot;) = &quot;Fratercula&quot;</code></pre>\n<p>BM25 full-text search was another obvious and much-demanded query plan. Similar\nto attribute search, full-text search works by first finding the documents that\nhave the query term present (commonly called &quot;postings&quot;). For an FTS index, we\nalso include the <code>(term count, document length)</code> metadata necessary for BM25\nscoring:</p>\n<pre><code>K::FTS(&quot;description&quot;, &quot;Atlantic&quot;) -&gt; vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]\nK::Attr(C0L0, &quot;description&quot;) -&gt; &quot;A sharply dressed black-and-white seabird with a \\\nhuge, multicolored bill, the Atlantic Puffin is often \\\ncalled the clown of the sea. It breeds in burrows on \\\nislands in the North Atlantic, and winters at sea.&quot;</code></pre>\n<p>Over time, we&#39;ve shipped several other index structures and query engines: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built around the same vector-primary storage layout.</p>\n<p>The ANN primary index has largely remained intact until today for one simple reason: it works really, really well for ANN search on object storage. On top of this architecture, we&#39;ve pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. Any significant change here risks introducing regressions in ANN performance.</p>\n<p>However, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization.</p>\n<p>As described above, turbopuffer currently puts the full contents of each document under its ANN address. When there is only one vector, the non-vector data is stored alongside the vector only once.</p>\n<p>However, for multi-vector representations of a document, such as document nesting or late interaction, this means we have to duplicate the contents for each vector. This is the reason for some of our more unfortunate limits.</p>\n<p>Any time a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to ensure they remain well clustered (otherwise recall may suffer). Because everything in a document is stored keyed by the ANN address of the document&#39;s vector, this rebalancing cascades to moving the full document contents, as well as any inverted (attribute and FTS) indexes that reference it. Updating just one vector can move hundreds of attributes and their indexes.</p>\n<p>This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.</p>\n<p>Modern query engines are vectorized: they run tight loops over blocks of values, which amortizes fixed per-block costs, compresses better, keeps the CPU pipeline full, and unlocks SIMD. DuckDB, for example, works in batches of 2,048 rows, ClickHouse up to ~65k, Lucene&#39;s posting blocks are 256 docs, and our ANN index works best with clusters of around 100–200 documents. Every query plan has an optimal block size, but today they are all constrained by the ANN primary index. A plan that wants blocks of thousands of documents to keep the CPU saturated is still stuck at 100–200.</p>\n<p>We&#39;ve already documented how much this matters in turbopuffer. Our first version of full-text search partitioned posting lists along ANN cluster boundaries, and the median block held just ~1.5 postings. FTS v2 reworked postings into fixed blocks of ~256, and the index got 10x smaller and queries got up to 20x faster. Posting lists could do that because they&#39;re stored separately and point at documents, so their layout doesn&#39;t have to follow the clusters. Aggregations and other scans read the documents themselves, and those are stored one block per cluster. As long as the ANN address is the primary key, their block size is constrained to the cluster size, even if they&#39;d prefer something larger.</p>\n<p>The solution to these problems is simple: don&#39;t key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.</p>\n<p>v3 is a new foundation that will unlock significant performance improvement on\nall query plans, and we hit a major milestone earlier this month: 100% of CI\npasses on turbopuffer v3. We started by focusing on correctness. Now we will\nmake it correct <em>and</em> fast. Watching benchmark numbers go down is great fun, so\nwe wanted to get you in at day zero of perf grinding. We will share\nthe benchmarks in public over the coming weeks, as we work toward (and\nbeyond) performance parity before rolling out v3 to production.</p>\n<p>turbopuffer is a fast search engine that hosts 1T+ documents, handles 10M+ writes/s, and serves 25k+ queries/s. We are ready for far more. We hope you&#39;ll trust us with your queries.</p>\n<p>Get started</p>","headings":[]}}