Skip to content

Staged Document Search

StrataSearch, a document search pipeline I built: when a filtered search finds too few matches, it drops a filter and searches again. This page rebuilds its offline search in your browser and measures what dropping filters costs.

Original built with Python · Voyage, Turbopuffer, OpenAI SDKs

What I found

3 of 3 filters

dropped one at a time while the search found fewer than 6 candidate notes, its default threshold, and it still reports “ready”.

2 of 4 outside

results fail a filter the request asked for; one is a 2022 lab experiment. StrataSearch’s output does not flag them: the red labels in the demo are this page’s addition.

2 of 4 returned, both match

at thresholds 1 to 3, where nothing is dropped and the search reports a shortfall instead.

Try it

Asked for 4 notes matching 3 filters and containing “cache”, the search drops filters while it has fewer than 6 candidates, its default threshold. Here it drops every filter, reports “ready” and returns 4 notes, two of which do not match the request. At thresholds 1 to 3 it drops nothing and reports a shortfall: 2 notes, both matching.

StrataSearch relaxes a query that comes back thin: when its first search finds fewer candidates than a threshold, it drops a filter and searches again. Each row runs the same request at a different threshold. Pick a row to see its results underneath: try the default, then 1 to 3.

The requestyear ≥ 2025kind = guidecollection = corecontains “cache” hard, never dropped· 4 results

matches every filter asked forfails one, let in by a dropped filter

ThresholdFilters droppedResults returnedMatch the request
none2 of 4 · shortfall2 of 2
year ≥ 2025kind = guide3 of 4 · shortfall2 of 3
year ≥ 2025kind = guidecollection = core4 of 4 · ready2 of 4

Filters go in a fixed order, year first, then kind or title, then topic or tags, then collection, until the first search finds at least as many candidates as the threshold.

Ready4 of 4 results at thresholds 6 or more

Filters dropped because the first search found fewer than 6 candidates: year ≥ 2025kind = guidecollection = core

NoteFieldsHard criteriaMatches the request
Search latency cacheguide · core · 2025yes
Search cache evidenceguide · core · 2025yes
Latency budget experimentexperiment · lab · 2022no, outside the request
Cache invalidation contractdesign · core · 2024no, outside the request

A struck value fails a filter you asked for. A hard criterion is a term whose every word must appear in the note. The “outside the request” flag is this page’s addition, checked against your original filters: StrataSearch’s own output does not mark these results.

Rejected by a hard criterion, and never used to pad a short list: Index recovery notebook, Retrieval channel fusion, Document tag migration, Filter relaxation trace.

Edit the query, see the scores, export a run

The pipeline, on StrataSearch’s eighteen example notes

Criteria are comma-separated. Every word of a hard criterion must appear in the note; soft criteria only rank notes, never exclude them.

Filters

Ready · 4 of 4 results at thresholds 6 or more

Channels
NoteFused scoreBody rankTitle, tags rankSoft criteria
Search latency cachenote-a0.0328113
Search cache evidencenote-p0.0323223
Latency budget experimentnote-e0.0317333
Cache invalidation contractnote-c0.0313440

The two channels are two searches, one over a note’s body and one over its title and tags; a rank is a note’s place in one of them. The fused score is reciprocal rank fusion, 1/(60 + rank) added over both channels. Soft criteria: how much of each one a note covers, 0 to 3.

Attempts

  1. 3 body hits under year ≥ 2025, kind = guide, collection = core
  2. 3 body hits under kind = guide, collection = core
  3. 5 body hits under collection = core
  4. 9 body hits under no filters
The eighteen notes
NoteKindTopicCollectionYearTags
Search latency cache note-aguideretrievalcore2025search, cache
Index recovery notebook note-bexperimentretrievallab2023search, recovery
Cache invalidation contract note-cdesignstoragecore2024cache, revision
Retrieval channel fusion note-dguideretrievalcore2025search, ranking
Latency budget experiment note-eexperimentretrievallab2022latency, cache
Storage replication protocol note-fdesignstoragecore2025storage, replication
Document tag migration note-gguideretrievalarchive2021search, schema
Filter relaxation trace note-hdesignretrievalcore2025filters, search
Vector namespace setup note-iguideretrievallab2024vectors, schema
Queue backpressure note-jdesignsystemscore2025queue, limits
Provider retry boundaries note-kguidesystemscore2024retry, providers
Ranking rubric inspection note-ldesignretrievallab2025ranking, evidence
Cache eviction sketch note-mexperimentstoragelab2023cache, memory
Replay artifact format note-nguidesystemscore2025replay, integrity
Tokenization limits note-oguideretrievalarchive2020tokens, limits
Search cache evidence note-pguideretrievalcore2025search, cache, latency
Scorer failure policy note-qdesignretrievalcore2024ranking, failure
API request inspection note-rguidesystemslab2025inspection, privacy
Contents

What does a search give up to fill its quota?

A search with filters can come back thin. StrataSearch's answer is relaxation: when its first search finds fewer candidates than a threshold, it drops one filter and searches again, until it has enough or has nothing left to drop. Its README is plain that filters are discovery hints, not boundaries, and that relaxation can erase useful constraints. This page makes the cost concrete: how many of the results a relaxed search returns are no longer what was asked for, and whether anything in the result says so.

What the audit found

StrataSearch's own example asks for year ≥ 2025, kind = guide, collection = core, the hard criterion “cache” (a term whose every word must appear in a note) and 4 results. At its default threshold of 6, the first search finds 3, 3, 5, 9 candidates as year, kind and collection drop in turn: every filter goes. The search reports ready. Two of its 4 results fall outside the request, among them a 2022 experiment from the lab collection, “Latency budget experiment”.

At thresholds 1 to 3 nothing is dropped: 2 results, both inside the request, and the search reports a shortfall. Neither status is wrong. “Ready” means the quota was met, and StrataSearch says so; it does not mean the quota was met with what was asked for. StrataSearch records which filters it dropped, but not which results only got in because of it. That per-result check is this page's addition, computed against the query's original filters.

The original system

StrataSearch is a document search pipeline I built in Python. A language model turns a request into filters; vector and keyword retrieval find candidates, relaxing the filters when results run thin; the result lists are fused, reranked and scored against the request's criteria. It is a neutralized public release of a staged private prototype. This page rebuilds its offline mode, the one that runs with no model and no network, and matches 6 of 6 recorded runs of it.

What this page is not

A measure of search quality. It runs on eighteen made-up notes with keyword matching only, and shows what relaxation does to a request, not how good the search is.

For engineers

The two channels and their fusion, the gate, the check against the original's CLI, and how to reproduce every number here.

How the pipeline runs

Two lexical channels. The body channel scores each unique query token by a damped term frequency, weighted by an inverse document frequency taken over the whole corpus and scaled by body length; the second channel adds 3 for a title match and 2 for a tag match, per token. Only positive scores enter a channel. Neither is an embedding or canonical BM25.

Relax, then reuse. The body channel runs first, dropping year, then kind or title, then topic or tags, then collection, until it has as many candidates as the threshold. The second channel then runs on the final, relaxed filters, and the two lists fuse by reciprocal rank, 1/(60 + rank) from each list a note appears in. Body only skips the second channel.

Gate, never pad. The reranker is an identity pass in the offline pipeline. A hard criterion passes when every one of its tokens appears in the note; soft criteria score token coverage from 0 to 3, rounded as Python rounds, halves to even. Only the first 50 candidates are judged. Eligible notes sort by hard passes, soft total, then fused score; a note that fails a hard criterion is rejected and never used to fill the quota.

Checked against the original

The code is a step-for-step port of local.py, retriever.py and pipeline.py, with the arithmetic in the source's order. It reproduces 6 of 6 recorded runs of the source's own CLI, among them body-only, two thresholds, a larger quota and a second query, matching every channel, score, selection and rejection to nine decimal places.

Where this comes from

StrataSearch's full pipeline is an LLM query planner that maps criteria to filters, vector and keyword retrieval with relaxation, rank fusion, a learned reranker, and an LLM criterion scorer, with local inspection and replay. Its live adapters for Voyage, Turbopuffer and OpenAI are real code, tested with injected clients; the release made no live provider calls, and the offline mode this page ports is the one that runs without any.

The source publishes no relevance or recall figures, only its fixture outcome, which this page reproduces: channel counts 3, 3, 5, 9, and note-a, note-p, note-e, note-c selected.

Limits in detail

Eighteen fictional notes and lexical features: no embeddings, no learned reranker, no LLM planner or judge, and no relevance judgments, so nothing here measures search quality. Tokens are ASCII and exact; there are no synonyms or stems. The corpus is the source's example and is not editable here; the query and settings are. An exported file is checked for consistency, not authenticity.

Reproduce it

Pick a row of the ladder, then open “Edit the query, see the scores, export a run”: drop the threshold to 3 and watch the shortfall; switch to body only; add a hard criterion that half the notes miss. Export a run and import it: the query reruns and the file is refused if its results do not follow. Files are limited to 50 KB. The port, the ladder and the tests are in the code for this page. From the site's Next.js app:

npx vitest run src/lib/projects/staged-search src/components/projects/staged-search