![]()
By Ming Ying on October 1, 2026

Two weeks ago, PlanetScale unveiled TIN, a full-text hunt expansion for Postgres. Their initiate article reported notable achievement wins complete a subset of ParadeDB’s content hunt functionality, specifically BM25-ranked content hunt and document counts.
We’d akin to broaden kudos to the PlanetScale squad 1. It’s awesome to see another Postgres phase investing in hunt (turns out group desire to hunt their relational data), and it’s apparent that a lot of thoughtful engineering went into TIN. We’re additionally blessed to see PlanetScale’s acceptance of ParadeDB’s benchmarker tool, which we built for exactly this benevolent of testing.
Let’s be extremely apparent concerning one thing: TIN is accelerated (at smallest 8x faster than ParadeDB 0.25 in all PlanetScale benchmark). So accelerated that the lone reply which made awareness was to close up and put on our achievement optimization hats. Two weeks later, here’s the BM25-ranked before and after, using the identical StackExchange benchmark dataset, harness, and device types (although TIN isn’t open-source so it’s operating on PlanetScale)2:
Queries
View
Percentile
↑ Higher is better
Warm-cache, read-only runs. TIN uses dense_ratio=2 alongside hunt elision disabled; see Handling Common Terms. One-second buckets; latency uses nearest-rank percentiles. Lines use a centered 9-second moving average; fable values are unsmoothed full-run results.
What’s engaging is not that we quickly closed the gap, but how we closed it. TIN’s article claims that their achievement is because of a essential architectural difference that uses Postgres’ inner ctid sectors as document identifiers. However, we closed this gap through a few optimization passes that had small to do alongside how documents are identified. We additionally tweaked several benchmark settings that didn’t provision a completely fair difference — additional on this later.
Let’s unpack our fixes and the configuration changes one by one.
An Overview of Text Search, and How TIN Claims to Be Faster
Without Text Index
With Text Index
The bosom of any content hunt indicator is a postings list: a per-term catalog of document identifiers containing that term. For instance, if an indicator has documents 1 to 10 and the term “database” appears in documents 2 and 4, the postings catalog for “database” is merely [2, 4]. Postings lists authorize you to acknowledge documents matching particular conditions extremely efficiently.
Tantivy, the hunt archive rearward ParadeDB, uses sequential u32 document IDs for its postings. These identifiers are inner to Tantivy, and are assigned based purely on insertion order. For the remainder of this post, DocId refers to the u32 document identifier used by Tantivy and ParadeDB.
Postgres identifies its rows by ctid values. A ctid is a tuple pointing to a row’s bodily location in Postgres’ block-based storage. (190, 17) identifies the row which is currently established in slot 17 of obstacle 190.
Because ParadeDB is a Postgres indicator powered by Tantivy, there have to be a map between DocId and ctid values. The crux of TIN’s article is that using ctid values immediately as document identifiers eliminates this map and enables productive bitmap operations and visibility checks. PlanetScale attributes much of TIN’s achievement advantage to the downstream benefits of that choice.
But Is It All About a Different Document Identifier?
The initiate article benchmarks two broad query types: Top K matches by BM25 mark and COUNT queries complete matching documents.
For counts, the ctid disagreement made sense. When millions of matches necessitate visibility checks, translating DocId values into ctid values adds up. Organizing postings about Postgres pages creates opportunities to peruse small data and lot that work.
For BM25 Top K queries, we were skeptical. ParadeDB defers ctid lookups until the final Top K documents have been gathered. For a top 10 query, that method 10 lookups. These lookups aren’t free, but they’re small in the overview and don’t explain an orders-of-magnitude gap.
Instead, we suspected that we could near the gap alongside assorted optimization opportunities elsewhere in our code.
This article focuses on our Top K BM25 optimizations. We’ve additionally optimized COUNT, which volition arrive in Part II.
Optimization 1: Reducing Random Access During BM25 Scoring
We started alongside a uncomplicated query: provision me the ten most applicable documents containing a sole term ordered by BM25 score. For faster local iteration, we used the smaller 28.7M Hacker News dataset.
EXPLAIN (ANALYZE, BUFFERS) SELECT id, title, by FROM hn_items WHERE heading === 'database' ORDER BY pdb.score(id) DESC LIMIT 10;
TIN touched far small Postgres pages than ParadeDB on this query, so we suspected that was the chief logic it was faster. This would additionally compound on the StackExchange dataset whenever several says arrive off disk. ParadeDB has a new characteristic that attributes leaf accesses to the data structures stored in those pages. It told us direct distant that we had an issue:
| Data structure | Share of leaf accesses |
|---|---|
| Fieldnorms | 1,513 (83%) |
| Everything alternatively (postings, metadata, etc.) | 311 (17%) |
“Fieldnorms” encode the dimension of a document’s indexed field, which BM25 uses to normalize scores. A fieldnorm in Tantivy is tiny: a document dimension quantized into a single-byte fieldnorm_id value. How could item small document for so many reads?
The issue was locality. Tantivy stores fieldnorms separately from postings, in an gathering indexed by DocId. Reading a term’s postings is sequential, but fetching the corresponding fieldnorms can jump all complete that array. With Tantivy’s customary memory-mapped storage, this layout is apt fine3 since all inhabitant fieldnorm is a cheap recollection lookup, but in Postgres it meant touching approximately 1,500 distinct fieldnorm pages for this query.
Our fix was to shop a fieldnorm gathering alongside all postings list, in the identical command as the postings’ DocId values. Scoring could afterward peruse fieldnorms sequentially alongside postings, eliminating the scattered lookups. After this change, fieldnorm accesses dropped from 1,500 pages to fair 30 (!).
Before: Shared fieldnorm array
Postings Fieldnorms "database": [DocId values] [fieldnorm IDs for all documents] "rust": [DocId values] ...
After: Fieldnorm gathering per term
Postings Fieldnorms "database": [DocId values] "database": [fieldnorm IDs] "rust": [DocId values] "rust": [fieldnorm IDs] ... ...
The tradeoff is storage, since a document’s fieldnorm is now repeated for all distinct term it contains. Fortunately, this doesn’t necessarily average multiplying fieldnorm retention by the figure of terms. In real-world corpora, most conditions have abbreviated postings lists and correspondingly small fieldnorm arrays. For instance, this alter grew the 28.7M HN indicator by concerning 9%.
Optimization 2: Choosing the Right Blockmax Pruning Algorithm
Breaking out fieldnorms delivered a huge speedup for queries alongside a small figure of terms, but we were motionless not satisfied alongside our achievement in disjunction queries alongside many terms. For instance, this query matches documents containing any of these terms:
EXPLAIN (ANALYZE, BUFFERS) SELECT id, title, by FROM hn_items WHERE content ||| 'rust arc copy recollection safety get checker ownership life rules' ORDER BY pdb.score(id) DESC LIMIT 10;
In this query, we noticed that equal although buffer says following the former optimization cut by approximately 80%, query times lone dropped by 5%, suggesting that the bottleneck in this case was algorithmic.
We profiled and discovered that most of the period was spent in item called the Blockmax WAND loop.
For context: Blockmax is the norm algorithm used by hunt engines to efficiently skip former chunks of postings whenever executing disjunction (e.g. termA OR termB) queries. There are two families of Blockmax: WAND and MAXSCORE. We won’t go into the intricacies of how they activity (there are many of fine specialized blogs for this), but at a elevated level:
- “Blockmax” comes from the fact that we can partition postings into blocks and for all obstacle precompute and shop the maximum imaginable mark that any term from this obstacle could contribute to the final BM25 score.
- A obstacle is skipped if its max mark cannot perchance attack the current Top K threshold. WAND and MAXSCORE are two distinct ways of doing this skipping.
The tradeoff between WAND and MAXSCORE is how much activity they expend deciding what to skip. WAND skips more, but spends additional CPU cycles to do so. MAXSCORE skips less, but incurs small overhead. When queries merge additional terms, WAND’s overhead grows and can outweigh the activity it skips.
Tantivy uses WAND. Lucene additionally used WAND until 2023, whenever they introduced MAXSCORE for certain queries. Today, Lucene dynamically chooses either WAND or MAXSCORE depending on the query shape.
We implemented a MAXSCORE way alongside a uncomplicated choice heuristic that uses MAXSCORE for disjunctions alongside at smallest three conditions and sufficiently compact postings and WAND for everything else. For the query complete containing 10 terms, we not lone brought p50 latency downward by ~6x and p95 by ~8x, we are now 2x faster vs. TIN on our 28.7M HN dataset:
Results
Terms
Match
View
Percentile
↑ Higher is better
Warm-cache, read-only runs. TIN uses dense_ratio=0.1 alongside hunt elision enabled; see Handling Common Terms. One-second buckets; latency uses nearest-rank percentiles. The equivalent selector does not use to one-term queries. Lines use a centered 9-second moving average; fable values are unsmoothed full-run results.
Benchmark Configuration Changes
PlanetScale’s benchmarks were constructed fairly, alongside the elimination of two anomalies that unintentionally preferred TIN: a ParadeDB syntax oversight and how TIN handles average terms.
ParadeDB Syntax
We noticed that the TIN benchmarks used ParadeDB’s query cord parser, which accepts Tantivy’s mini query tongue via the @@@ operator. The issue is that these queries weren’t qualified alongside a site name, e.g. <query> alternatively of <field>:<query>.
When queries are unqualified, ParadeDB searches complete all indexed content sectors by default. In the StackExchange dataset, the two the id and build columns were indexed, which method ParadeDB was disadvantaged since it was searching complete two columns per query whereas TIN searched lone one.
To defender against this, we moved all ParadeDB queries to use our native ||| (disjunction), &&& (conjunction), and ### (phrase) operators.
Handling Common Terms
We were effortlessly beating TIN on the BM25 hunt queries in our HN benchmark. But whenever we loaded PlanetScale’s StackExchange dataset and queries, we were motionless 30% rearward on throughput since of our much longer rear latencies. How could we be multiple times faster on our benchmark but slower on PlanetScale’s?
It turns out the gap came from a scoring shortcut for average conditions that TIN calls dense-term elision, and we think its use complete the StackExchange dataset specifically is debatable.
A concise explainer: average conditions akin "the" and "is" have enormous postings lists that are costly to peruse and score. Yet BM25 weights them so low that they barely move the final ranking. Most hunt engines grip this at indexing period alongside a stopword dictionary (both engines assistance this, but it wasn't enabled in the benchmark). TIN takes a distinct approach: at query time, it skips scoring for any term that appears in additional than 10% of the corpus (configurable via dense_ratio). It's the most engaging idea in TIN from a hunt practitioner's item of view, and we volition expend additional period thinking concerning this.
Of way there is continually a tradeoff, and current it is correctness. With elision on, TIN computes an approximation of BM25 by ignoring average words, causing results to possibly arrive rear in a distinct command than true BM25 would produce. This normally isn’t a issue for most real-world queries, unless the query is made up entirely of average words.
When we looked at what the Stack Overflow benchmark ran, we were amazed to discover complete queries made up of these average words. That’s since they were generated by sampling successive term spans from the Stack Exchange corpus, which produced queries akin "is it", "to a", and "is to" 4. When we classified the largest latency gaps between ParadeDB and TIN, those identical queries (which are not genuine hunt queries) dominated the list. On them, TIN skipped most of the scoring activity during ParadeDB computed exact scores.
For TIN, alongside elision enabled, we established that 5:
- 47.8% of queries returned at smallest one outcome in the Top 10 that was not in the “true” Top 10.
- 6.4% of queries returned results anywhere none of the Top 10 were in the “true” Top 10 — in another words, all the results were wrong.
- Disjunctions were particularly inaccurate: at smallest one non-Top-10 outcome appeared in the Top 10 for 88.2% of disjunction queries.
| TIN Query Style | At smallest one outcome exterior the “true” Top 10 |
|---|---|
| Conjunction | 39.9% |
| Disjunction | 88.2% |
| Phrase | 13.8% |
| Overall | 47.8% |
To be clear, we’re not saying that "different from exact BM25" automatically method "worse". Elided conditions have low importance by design, and determining whether the elided results are small applicable would necessitate individual relevance judgments, which this benchmark does not have. But the workload is framed as BM25 top-K search, and alongside elision enabled, TIN and ParadeDB are not computing the identical ranking.
For this logic our header difference uses exact BM25 for the two engines, alongside TIN configured alongside dense_ratio=2. Below we additionally display TIN's default elision-enabled configuration, anywhere they “beat” us, since that is what the first benchmark used.
We’re sharing the two sets of results so readers can diagram their own conclusions. We don’t desire benchmark settings to distract from the achievement improvements we made to ParadeDB. At the identical time, we would be remiss not to citation them, as they create specified a big difference to the initially published results. For fine measure we have additionally included ParadeDB alongside stopwords enabled.
View
Percentile
↑ Higher is better
Warm-cache, read-only mixed-query runs. One-second buckets; latency uses nearest-rank percentiles. Lines use a centered 9-second moving average; fable values are unsmoothed full-run results.
So Is There a Superior Document Identifier?
As for the choice of TIN’s ctid vs. ParadeDB’s u32 document identifiers, we see it as a tradeoff as well.
A premise of the TIN article is that ctid values are a universally fine choice for postings lists. The issue alongside this is that nothing compresses improved than dense, sorted, distinctive integers. Consequently, most hunt systems use u32 DocId values. Switching to a 48-bit document identifier is not inherently additional efficient, particularly since the 48 bits in a ctid are the chain of numbers from two distinct domains (block numbers, numbered in the millions, and tuple offsets, at most 291).
Dense u32 DocId values have another advantage: they create it uncomplicated to nexus postings to columnar storage. Postings inform you which documents match; columns let you efficiently admission those documents’ metadata attributes (like numeric values or category labels).
Column stores are frequently connected alongside OLAP databases, but they’re additionally crucial for hunt queries:
- Top K by field: “Give me products matching my query, ordered by price.”
- Range filters: “Give me matching products between $50 and $100.”
- Faceting: “Give me the top 10 products, and aspect by the figure of matches in all category.”
All of these hunt queries necessitate a columnar format. With DocId values, the association is straightforward. Within a segment, document 42 corresponds to row 42 in all column. Once a postings catalog gives us that ID, we can appearance up its columnar value directly.
A ctid doesn’t provision us that pillar position. It identifies a bodily Postgres location, specified as leaf 190, slot 17. To recover that document’s cost from a column, we archetypal need to decide which pillar row corresponds to (190, 17). That requires a mapping or an equal lookup.
TIN is awesome at BM25 scoring and document counting, but that’s fair the tip of what makes up a hunt motor akin Elasticsearch. For the “rest of search”, you need a columnar representation. If TIN decides to create one, we doubtful they’ll have to pay the identical ctid/DocId translation disbursal (but in the reverse direction).
Why Tantivy Remains the Right Choice for Us
TIN’s article attributes its achievement advantage to using ctid values alternatively of DocId values. But we closed the gap without changing our document identifiers by:
- Writing denser data structures alongside improved locality.
- Using a Blockmax algorithm alongside higher throughput.
- A few another small optimizations connected to lazy study of another pieces of data.
Most of these changes happened in Tantivy, our hunt library.
Over the former few years we’ve frequently debated whether construction on Tantivy was the correct choice for ParadeDB, versus what appears to be the TIN method of penning a new hunt motor from scratch. This inquiry has reinforced our condemnation in our decision to use Tantivy. It brings complete a decade of development, battle evaluation by several of the world’s largest companies, and notable speed. Tantivy doesn’t continually mesh absolutely alongside Postgres’ obstacle layout correct off the shelf, but its extensibility and features additional than create up for that.
You power ask: why hadn’t we done these optimizations already? Performance activity never ends, and our engineering resources are finite. After reaching Elasticsearch parity on our content hunt benchmarks, we shifted our notice toward expanding ParadeDB’s capabilities beyond “just content search” into efficiently executing hunt queries that affect complex filters, facets, and joins. By now our hunt API is extremely broad, so we’re glad that TIN brought our notice to this chance for center optimization.
Closing Thoughts
The open origin hunt community has a longstanding tradition of collaboration. For instance, notwithstanding being rivalrous hunt libraries, Lucene and Tantivy frequently portion ideas and benchmark against all another in a affable way. We explored how this benefits the two projects in our conversation alongside Paul Masurel, creator of Tantivy. We anticipation this becomes another example of this. It’s been a fun sprint for us.
We value that the TIN authors shared several of their engineering decisions in their blog, equal if the project itself isn’t open source. All of our activity is open, and we’ve already begun to upstream the applicable improvements from this spike to Tantivy.
We’ve cut a 0.26.0-rc.2 publish applicant so these results are reproducible. For existing ParadeDB users, these improvements volition be folded into the next stable release, 0.26.0, targeted for next week. We’ve made certain that these changes are backwards compatible, although a reindex volition be required to inherit all the optimizations.
To our community contributors: this inquiry was time-boxed, and we think there are many additional optimization strings to drag on (especially in the direction of additional productive Blockmax pruning and reducing buffer access). We greeted any contributions that shove the achievement frontier further.
In the next part we’ll conversation the optimizations we made about our COUNT achievement (hint: they additionally didn’t necessitate changing our document identifiers). Until then, blessed searching! We’re enthusiastic to provision faster content queries to the Postgres and hunt communities.