RIP, vector database

Hacker News by 7 min read 510x views
RIP, vector database

Share Post

September 30, 2026•Dan Harrison (Engineer)

We are changing turbopuffer's retention architecture. The new engine, informally called "turbopuffer v3", changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It volition authorize us to assist many more query plans, at much greater scale.

turbopuffer launched as a serverless vector database, highly specialized to the project of serving extremely cheap and fairly fast vector searches. Object retention as the origin of fact gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion.

Over time, turbopuffer has rotate into a generalized repository for hunt (and other things). The query motor has evolved alongside the way to assistance all of these query plans, but the retention architecture has remained mostly unchanged: the ANN vector indicator was and motionless is the chief indicator about which all another indexes and query plans revolve.

We've pushed the chief ANN indicator as far as we can, but it's period to move on. We're in the procedure of moving to a new chief index, and making ANN "just another" secondary index. We idea it power be fun to open up the doors and let you prosecute along.

For this archetypal update, we'll set the phase alongside why we're doing this in the first place. Walk alongside me on a abbreviated journey from tpuf v1 to today.

v1: an ID and a vector

In the archetypal type of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing knowledge at the period was graph-based vector indexes, but a hierarchical clustering indicator plays improved alongside entity storage. We started with SPANN, and eventually migrated to SPFresh to assistance incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree alongside a sole root.

 ┌───────────────────┐ │ base centroid │ └───────────────────┘ ╱ │ ╲ ╱ │ ╲ ┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐ │ leaf centroid │ │ leaf centroid │ │ leaf centroid │ └───────────────────┘ └───────────────────┘ └───────────────────┘ ╱ ╲ ╱ ╲ ╱ ╲ ╱ ╲ ╱ ╲ ╱ ╲ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ │ vector │ │ vector │ │ vector │ │ vector │ │ vector │ │ vector │ └────────┘ └────────┘ └────────┘ └────────┘ └────────┘ └────────┘ 
 ┌───────────────┐ │ base centroid │ └───────────────┘ ╱ │ ╲ ╱ │ ╲ ┌────────┐┌────────┐┌────────┐ │ leaf ││ leaf ││ leaf │ │centroid││centroid││centroid│ └────────┘└────────┘└────────┘ ╱ ╲ ╱ ╲ ╱ ╲ ┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐ │vec││vec││vec││vec││vec││vec│ └───┘└───┘└───┘└───┘└───┘└───┘ 

We implemented this on top of a retention tier presenting as a key-value map, with sorted and distinctive keys. Each collection is stated a ClusterId, and vectors within all collection are stated a compact LocalId.

 K::Vector(C0L0) = vec![0.45, 0.32, ...] K::Id(C0L0) = 7 K::Vector(C0L1) = vec![-0.28, 0.96, ...] K::Id(C0L1) = 13 K::Vector(C1L4) = vec![0.64, -0.48, ...] K::Id(C1L4) = C0 

As you can see above, everything is keyed by ClusterId and LocalId (e.g. C0L1), which together we call the ANN address. This is what we average when we say the ANN indicator is the chief index.

Two new query plans marked the informal passage from turbopuffer v1 → v2: attribute filtering and full-text search.

Attribute filtering

Naturally, customers wanted to be capable to add trait values and display vector searches on them. To make filtering accelerated and high-recall, we modeled these as an inverted indicator that maps an trait value to the ANN location of the documents that merge it.

K::AttrIndex("family", "Alcidae") -> vec![C0L3, C1L2, C1L3, ...] K::AttrIndex("genus", "Fratercula") -> vec![C0L3, C1L2, C1L9, ...] 

For projections (include_attributes), we additionally stored the document attributes alongside the ID and the vector.

K::Vector(C0L0) = vec![0.45, 0.32, ...] K::Id(C0L0) = 7 K::Attr(C0L0, "family") = "Alcidae" K::Attr(C0L0, "genus") = "Fratercula" 

BM25 full-text hunt was another apparent and much-demanded query plan. Similar to trait search, full-text hunt plant by archetypal finding the documents that have the query term current (commonly called "postings"). For an FTS index, we also contain the (term count, document length) metadata necessary for BM25 scoring:

K::FTS("description", "Atlantic") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...] K::Attr(C0L0, "description") -> "A sharply dressed black-and-white seabird alongside a \ huge, multicolored bill, the Atlantic Puffin is frequently \ called the clown of the sea. It breeds in burrows on \ islands in the North Atlantic, and winters at sea." 

Over time, we've shipped multiple another indicator structures and query engines: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built about the same vector-primary retention layout.

The issue alongside a vector chief index

The ANN chief indicator has mostly remained intact until today for one simple reason: it plant really, really fine for ANN hunt on entity storage. On top of this architecture, we've pushed vector hunt to single indexes of 100B+ vectors serving 200 ms p99 says at 1k+ QPS. Any important alter current risks introducing regressions in ANN performance.

However, this layout holds us rear from being state-of-the-art for the non-vector query shapes we support, in three chief ways: retention amplification, write amplification, and constricted vectorization.

Storage amplification

As described above, turbopuffer currently puts the complete contents of each document under its ANN address. When there is lone one vector, the non-vector data is stored alongside the vector lone once.

However, for multi-vector representations of a document, specified as document nesting or delayed interaction, this method we have to copy the contents for each vector. This is the logic for several of our additional unfortunate limits.

Write amplification

Any period a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to justify they remain fine clustered (otherwise recall may suffer). Because everything in a document is stored keyed by the ANN location of the document's vector, this rebalancing cascades to moving the complete document contents, as fine as any inverted (attribute and FTS) indexes that citation it. Updating fair one vector can move hundreds of attributes and their indexes.

This compose amplification is ample adequate that our efforts to tune indexing throughput have started to hit diminishing returns.

Limited vectorization

Modern query engines are vectorized: they run tight loops complete blocks of values, which amortizes fixed per-block costs, compresses better, keeps the CPU pipeline full, and unlocks SIMD. DuckDB, for example, works in batches of 2,048 rows, ClickHouse up to ~65k, Lucene's posting blocks are 256 docs, and our ANN indicator plant finest alongside clusters of about 100–200 documents. Every query scheme has an optimal obstacle size, but today they are all constrained by the ANN chief index. A scheme that wants blocks of thousands of documents to keep the CPU soaked is motionless stuck at 100–200.

We've already documented how much this matters in turbopuffer. Our archetypal version of full-text hunt partitioned posting lists alongside ANN collection boundaries, and the median obstacle held fair ~1.5 postings. FTS v2 reworked postings into fixed blocks of ~256, and the index got 10x smaller and queries got up to 20x faster. Posting lists could do that since they're stored separately and item at documents, so their layout doesn't have to prosecute the clusters. Aggregations and another scans peruse the documents themselves, and those are stored one obstacle per cluster. As lengthy as the ANN location is the chief key, their obstacle size is constrained to the cluster size, equal if they'd favor item larger.

RIP, chief vector index

The resolution to these problems is simple: don't key on the ANN address. That is precisely the alter turbopuffer v3 makes. As you can imagine, it is not a trivial change.

We lately hit a important milestone: As of before this month, 100% of CI passes on turbopuffer v3. However, it represented a important achievement regression from manufacturing turbopuffer. That's not a surprise! So far, we've focused primarily on the foundational scheme and correctness, and we have fair started tuning it. Watching the numbers go downward is awesome fun, so we wanted to get you in at day zero of perf grinding.

We'll add results to the benchmarks as we compose up what changed. Along the way, we'll dive into item concerning the new architecture and the optimizations that volition ultimately create turbopuffer faster for additional query plans in prod.

turbopuffer

turbopuffer is a accelerated hunt motor that hosts 1T+ documents, handles 10M+ writes/s, and serves 25k+ queries/s. We are prepared for far more. We anticipation you'll rely us alongside your queries.

Get started

Other Article Hacker News
↑
Close Right Ads
Close Left Ads