SEEKERbuilding a search engine from first principles
Playground
OverviewSearch playgroundThe corpusDifferential tests
I · Foundations
1What a Search Engine Actually Does2Documents, Fields & Data Types3Tokenization, Normalization & Stemming
II · The First Search Engine
4Inverted Index5Postings & Positions6Query Execution
III · Relevance
7Ranking8BM25 from First Principles9Unknown Words, OOV Terms & Zero Results
IV · Query Language
10Text vs Keyword11Match, Term, Bool & Range12Phrase, Prefix, Wildcard & Regex
V · Fuzzy Search
13Edit Distance14Levenshtein Automata15Candidate Terms & Expansion
VI · Term Dictionaries
16Tries17Finite-State Machines18FSTs19BlockTree
VII · Numeric & Spatial Search
20Points & Range Indexing21KD Trees22BKD Trees
VIII · Storage Engine
23Segments24Byte-Level Formats25Compression & Checksums26mmap & Page Cache
IX · Production Query Engine
27Query Planning28Caching29Concurrency & Thread Pools30Painless & Scripted Scoring
X · Distributed Search
31Shards32Replicas33Distributed Query & Fetch34Routing & Rescoring
XI · Cluster & Operations
35Refresh, Flush & Commit36Recovery & Replication37Cluster Management
XII · Production Engineering
38Performance Engineering39Capacity Planning40Observability & Failure Modes
XIII · The Seeker Project
41Implementation Roadmap42Testing Strategy43From Seeker to OpenSearch
Part 2 · The First Search Engine·Chapter 4

Inverted Index

term → set(docID). One map turns scanning every document into a single lookup.

The question
Which documents contain this term?
The structure responsible
Inverted index

In this lab: Build the index a document at a time, then race it against a full scan.

building the index…
← Chapter 3
Tokenization, Normalization & Stemming
Chapter 5 →
Postings & Positions