The company behind turbopuffer, a serverless vector database that launched years ago as a specialized search tool, has announced that it is rebuilding its entire core product from scratch. It is telling everyone about it. The change, which the company calls “turbopuffer v3,” will make search faster across the board — including vector search — and set the stage for quicker SQL queries as well. The company is not subtle about what it is doing away with; it calls the old design by a blunt name: RIP, vector database.
This disclosure stands apart from what most firms do. Typically, companies guard their internal architecture closely, particularly during a rewrite of the foundations. But turbopuffer has chosen to open the doors and let followers observe the process as it happens. The initial update lays out the groundwork by tracing the product’s history, moving from its first version right up to the present day.
From ID and Vector to Everything Else
When turbopuffer launched as version 1, documents were simple: an ID and a vector. The vector represented meaning, and the ID pointed to it. The storage architecture was built around a hierarchical clustering index, starting with SPANN and later migrating to SPFresh to support incremental indexing. Vectors were grouped into clusters, whose centroids were clustered again, forming a tree with a single root.
In the announcement, the company laid out the full design, with the root centroid perched at the summit and descending branches reaching through levels of leaf centroids, each one directing toward its own vectors. Each grouping receives a ClusterId, and each item inside a group receives a LocalId. Joined together, these two labels form what the company refers to as the ANN address — the main identifier for the entire arrangement.
Vector search over object storage operated smoothly at first, but the system reached its boundaries quickly. Once clients began requesting the ability to add attribute values and filter searches based on those attributes, turbopuffer found itself needing to construct an inverted index. This index connects an attribute value with the ANN addresses of the documents that hold it. The company presented several examples of how this works.
K::AttrIndex("family", "Alcidae") -> vec![C0L3, C1L2, C1L3, ...]
K::AttrIndex("genus", "Fratercula") -> vec![C0L3, C1L2, C1L9, ...]
The postings themselves became the starting point for full-text search rather than the documents they pointed to. A query would locate the postings containing its term, and then metadata was added alongside those postings to enable BM25 scoring, including term counts and document lengths. Here is another illustration:
The call K::FTS(“description”, “Atlantic”) produces a vector with pairs, each containing an identifier such as C0L0 or C9L4 followed by two values, including 2, 37, 1, and 42. The second line, K::Attr(C0L0, “description”), yields a single result of “A sharply dressed black-and-white seabird with a huge, multicolored bill, the Atlantic Puffin is often called the clown of the sea. It breeds in burrows on islands in the North Atlantic, and winters at sea.”.
Why the Old Design Held On
The primary index built around the ANN design proved highly effective for ANN search on object storage. The company has since expanded that architecture to support single indexes containing 100 billion-plus vectors, delivering 200 millisecond p99 reads while handling 1,000-plus queries per second. Because of this, any major alteration carries a real risk of degrading ANN performance.
The design kept the firm behind the curve when it came to the non-vector query shapes it supports. Three major issues stand out: storage expansion, write expansion, and limited vectorization.
The entire contents of each document sit beneath its ANN address, which is what causes storage amplification. With just one vector, the non-vector data gets stored only once, alongside it. But when a document carries multiple vectors — through document nesting or late interaction — the company must duplicate those contents for each vector instead. This duplication explains some of its more regrettable constraints.
SPFresh can rebalance vectors whenever a document is added, changed, or removed, to keep them well clustered. Since all document data is stored using the ANN address of the document’s vector as the key, that rebalancing moves the entire document contents along with any inverted (attribute and FTS) indexes that point to it.
The New Primary Index
The firm is shifting its main index to a fresh one while keeping ANN “just another” as a supporting index. This arrangement has limited how certain query plans work, such as GROUP BY and aggregations. The firm has carried the vector-primary design as far as possible, and now it is ready for change.
This change speeds up search across the board, including vector search, for turbopuffer. It also sets the groundwork to shift far more SQL queries over to turbopuffer and run them at full speed. The announcement presents the move as a logical step forward from turbopuffer v1 to v2, using the addition of two new query plans — attribute filtering and full-text search — to mark the passage.
Turbopuffer powers the company’s earliest customers, including Cursor and Notion, validated the value of these tradeoffs. Linear’s syncing engine for tasks that do not involve searching, while the query engine has grown to accommodate all of these query plans. The storage design, however, has stayed much the same.
What the Company Is Building Next
The firm is shifting over to its new main index now. The opening statement explains the reasoning behind the move entirely. The enterprise considers it a pleasant prospect to open the doors and allow followers to tag along.
The comparison below shows the two architectures side by side.
| Architecture | Primary Index | Key Limitation |
|---|---|---|
| Old (vector-primary) | ANN vector index | Storage amplification, write amplification, limited vectorization |
| New (unspecified) | Unspecified primary | Not stated |
Since the beginning, the company has released a number of additional indexing systems and query tools, including aggregations, regular expression search, approximate matching, sparse vector search, and attribute sorting — all constructed with the same vector-first storage design at their core. The ANN primary index has stayed largely unchanged until now for one simple reason: it performs really, really well for ANN search on object storage.
Our View of the Announcement
The firm has laid bare the expenses tied to its former design and the grounds for moving away from it. It has also been unusually open about the limitations that shaped its choice.
The firm is banking on the new architecture to deliver quicker SQL queries without sacrificing the pace of vector search. It is a large wager. The company admits the danger of regression in ANN performance, a signal that the code overhaul is far from simple.
The company’s customers are already using turbopuffer for non-search purposes, including Linear’s syncing engine. The company has shipped several other index structures and query engines since the early days: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built around the same vector-primary storage layout.
The statement shows the firm looking past its existing product limits. The fresh main index will handle far more SQL queries than before.
The system has been adopted by early users like Cursor and Notion, and it is currently being used by Linear to power its syncing engine. The current ANN performance stands at 100 billion-plus vectors, with read times sitting at 200 ms p99 and supporting a throughput of 1,000+ queries per second. Its query capabilities include vector search, attribute filtering, full-text search, aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering.
This enterprise sits amid a shift. The initial report explains the reasoning behind the change entirely. The firm considers it pleasant to open its doors and permit followers to join in. The firm has driven its vector-primary design to its limits, and now it is time to proceed further.
Hard Numbers
- Old design: ANN vector index as primary
- Current ANN scale: Single indexes supporting 100 billion-plus vectors
- Read latency: 200 ms p99
- Throughput: 1,000+ queries per second
- New design: Unspecified primary with ANN as secondary
- Query capabilities: Vector search, attribute filtering, full-text search, aggregations, regex search, fuzzy matching, sparse vector search, attribute ordering
Source material: “RIP, vector database,” turbopuffer.com.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

