Full-text search is now up to 714× faster.
We built a new full-text index for Oliver. It finds a common word in 7 microseconds, takes up to 3.5 times less space, builds nearly five times faster, and returns exactly the same answers as our existing index.
Why we rebuilt it
Searching logs for a word is the question observability users ask all day: an error code, a customer ID, the word "timeout". Oliver has always indexed every word, so those searches never have to read every line. That index was also the heaviest thing Oliver built. On long log entries, building it was most of the work of getting new data ready to query.
So we wrote a new one. It keeps a dictionary of every word, and for each word a compact list of the rows it appears in and where in each row. Finding a word is one dictionary lookup and one short list.
Search by search
We measured both indexes on the same machine and the same data: a million short log lines, and a hundred thousand long entries of about 2 KB each.
A common word in long log entries went from 5 milliseconds to 7 microseconds, 714 times faster. In short log lines the same search is 144 times faster. Boolean and prefix searches moved from milliseconds to tenths of a millisecond, and the first search on data nobody has touched yet drops from tens of milliseconds to a few.
The smaller wins add up too. Looking up a single ID is 3.5 times faster, typo-tolerant search is 3.4 to 3.7 times faster, and phrases made of common words run 1.5 to 3 times faster.
Less to store and less to build
On long log entries the new index is 3.5 times smaller, and on short lines 2.3 times smaller. It builds 4.6 to 4.9 times faster, so new data becomes searchable sooner, and compaction uses about the same memory as before.
The same answers
A faster index is only worth having if it finds the same rows. We ran both indexes side by side on the same tables and compared every answer: 1,635 queries and 11,459 individual index lookups, before and after compaction, covering single words, AND, OR and NOT, prefixes, typo-tolerant search, regular expressions and phrases. They matched every time.
Turning it on
The new index is called Tokens, and it is available today in OliverDB Cloud, in-VPC deployments and Oliver for Snowflake. Choose it for any text column in the console's schema builder, make it the default for a whole deployment with one setting, or declare it in Terraform.
Our existing index stays the default, so nothing changes until you choose Tokens. When you do, existing columns move over as their data is compacted.