Staged Document Search
StrataSearch, a document search pipeline I built: when a filtered search finds too few matches, it drops a filter and searches again. This page rebuilds its offline search in your browser and measures what dropping filters costs.
Original built with Python · Voyage, Turbopuffer, OpenAI SDKs
What I found
dropped one at a time while the search found fewer than 6 candidate notes, its default threshold, and it still reports “ready”.
results fail a filter the request asked for; one is a 2022 lab experiment. StrataSearch’s output does not flag them: the red labels in the demo are this page’s addition.
at thresholds 1 to 3, where nothing is dropped and the search reports a shortfall instead.
Try it
Asked for 4 notes matching 3 filters and containing “cache”, the search drops filters while it has fewer than 6 candidates, its default threshold. Here it drops every filter, reports “ready” and returns 4 notes, two of which do not match the request. At thresholds 1 to 3 it drops nothing and reports a shortfall: 2 notes, both matching.
StrataSearch relaxes a query that comes back thin: when its first search finds fewer candidates than a threshold, it drops a filter and searches again. Each row runs the same request at a different threshold. Pick a row to see its results underneath: try the default, then 1 to 3.
The requestyear ≥ 2025kind = guidecollection = corecontains “cache” hard, never dropped· 4 results
matches every filter asked forfails one, let in by a dropped filter
| Threshold | Filters dropped | Results returned | Match the request |
|---|---|---|---|
| none | 2 of 4 · shortfall | 2 of 2 | |
| 3 of 4 · shortfall | 2 of 3 | ||
| 4 of 4 · ready | 2 of 4 |
Filters go in a fixed order, year first, then kind or title, then topic or tags, then collection, until the first search finds at least as many candidates as the threshold.
Ready4 of 4 results at thresholds 6 or more
Filters dropped because the first search found fewer than 6 candidates: year ≥ 2025kind = guidecollection = core
| Note | Fields | Hard criteria | Matches the request |
|---|---|---|---|
| Search latency cache | guide · core · 2025 | yes | |
| Search cache evidence | guide · core · 2025 | yes | |
| Latency budget experiment | experiment · lab · 2022 | no, outside the request | |
| Cache invalidation contract | design · core · 2024 | no, outside the request |
A struck value fails a filter you asked for. A hard criterion is a term whose every word must appear in the note. The “outside the request” flag is this page’s addition, checked against your original filters: StrataSearch’s own output does not mark these results.
Rejected by a hard criterion, and never used to pad a short list: Index recovery notebook, Retrieval channel fusion, Document tag migration, Filter relaxation trace.
Edit the query, see the scores, export a run
| Note | Fused score | Body rank | Title, tags rank | Soft criteria |
|---|---|---|---|---|
| Search latency cachenote-a | 0.0328 | 1 | 1 | 3 |
| Search cache evidencenote-p | 0.0323 | 2 | 2 | 3 |
| Latency budget experimentnote-e | 0.0317 | 3 | 3 | 3 |
| Cache invalidation contractnote-c | 0.0313 | 4 | 4 | 0 |
The two channels are two searches, one over a note’s body and one over its title and tags; a rank is a note’s place in one of them. The fused score is reciprocal rank fusion, 1/(60 + rank) added over both channels. Soft criteria: how much of each one a note covers, 0 to 3.
Attempts
- 3 body hits under year ≥ 2025, kind = guide, collection = core
- 3 body hits under kind = guide, collection = core
- 5 body hits under collection = core
- 9 body hits under no filters
The eighteen notes
| Note | Kind | Topic | Collection | Year | Tags |
|---|---|---|---|---|---|
| Search latency cache note-a | guide | retrieval | core | 2025 | search, cache |
| Index recovery notebook note-b | experiment | retrieval | lab | 2023 | search, recovery |
| Cache invalidation contract note-c | design | storage | core | 2024 | cache, revision |
| Retrieval channel fusion note-d | guide | retrieval | core | 2025 | search, ranking |
| Latency budget experiment note-e | experiment | retrieval | lab | 2022 | latency, cache |
| Storage replication protocol note-f | design | storage | core | 2025 | storage, replication |
| Document tag migration note-g | guide | retrieval | archive | 2021 | search, schema |
| Filter relaxation trace note-h | design | retrieval | core | 2025 | filters, search |
| Vector namespace setup note-i | guide | retrieval | lab | 2024 | vectors, schema |
| Queue backpressure note-j | design | systems | core | 2025 | queue, limits |
| Provider retry boundaries note-k | guide | systems | core | 2024 | retry, providers |
| Ranking rubric inspection note-l | design | retrieval | lab | 2025 | ranking, evidence |
| Cache eviction sketch note-m | experiment | storage | lab | 2023 | cache, memory |
| Replay artifact format note-n | guide | systems | core | 2025 | replay, integrity |
| Tokenization limits note-o | guide | retrieval | archive | 2020 | tokens, limits |
| Search cache evidence note-p | guide | retrieval | core | 2025 | search, cache, latency |
| Scorer failure policy note-q | design | retrieval | core | 2024 | ranking, failure |
| API request inspection note-r | guide | systems | lab | 2025 | inspection, privacy |
Contents
What does a search give up to fill its quota?
A search with filters can come back thin. StrataSearch's answer is relaxation: when its first search finds fewer candidates than a threshold, it drops one filter and searches again, until it has enough or has nothing left to drop. Its README is plain that filters are discovery hints, not boundaries, and that relaxation can erase useful constraints. This page makes the cost concrete: how many of the results a relaxed search returns are no longer what was asked for, and whether anything in the result says so.
What the audit found
StrataSearch's own example asks for year ≥ 2025, kind = guide, collection = core, the hard criterion “cache” (a term whose every word must appear in a note) and 4 results. At its default threshold of 6, the first search finds 3, 3, 5, 9 candidates as year, kind and collection drop in turn: every filter goes. The search reports ready. Two of its 4 results fall outside the request, among them a 2022 experiment from the lab collection, “Latency budget experiment”.
At thresholds 1 to 3 nothing is dropped: 2 results, both inside the request, and the search reports a shortfall. Neither status is wrong. “Ready” means the quota was met, and StrataSearch says so; it does not mean the quota was met with what was asked for. StrataSearch records which filters it dropped, but not which results only got in because of it. That per-result check is this page's addition, computed against the query's original filters.
The original system
StrataSearch is a document search pipeline I built in Python. A language model turns a request into filters; vector and keyword retrieval find candidates, relaxing the filters when results run thin; the result lists are fused, reranked and scored against the request's criteria. It is a neutralized public release of a staged private prototype. This page rebuilds its offline mode, the one that runs with no model and no network, and matches 6 of 6 recorded runs of it.
What this page is not
A measure of search quality. It runs on eighteen made-up notes with keyword matching only, and shows what relaxation does to a request, not how good the search is.
For engineers
The two channels and their fusion, the gate, the check against the original's CLI, and how to reproduce every number here.
How the pipeline runs
Two lexical channels. The body channel scores each unique query token by a damped term frequency, weighted by an inverse document frequency taken over the whole corpus and scaled by body length; the second channel adds 3 for a title match and 2 for a tag match, per token. Only positive scores enter a channel. Neither is an embedding or canonical BM25.
Relax, then reuse. The body channel runs first, dropping year, then kind or title, then topic or tags, then collection, until it has as many candidates as the threshold. The second channel then runs on the final, relaxed filters, and the two lists fuse by reciprocal rank, 1/(60 + rank) from each list a note appears in. Body only skips the second channel.
Gate, never pad. The reranker is an identity pass in the offline pipeline. A hard criterion passes when every one of its tokens appears in the note; soft criteria score token coverage from 0 to 3, rounded as Python rounds, halves to even. Only the first 50 candidates are judged. Eligible notes sort by hard passes, soft total, then fused score; a note that fails a hard criterion is rejected and never used to fill the quota.
Checked against the original
The code is a step-for-step port of local.py, retriever.py and pipeline.py, with the arithmetic in the source's order. It reproduces 6 of 6 recorded runs of the source's own CLI, among them body-only, two thresholds, a larger quota and a second query, matching every channel, score, selection and rejection to nine decimal places.
Where this comes from
StrataSearch's full pipeline is an LLM query planner that maps criteria to filters, vector and keyword retrieval with relaxation, rank fusion, a learned reranker, and an LLM criterion scorer, with local inspection and replay. Its live adapters for Voyage, Turbopuffer and OpenAI are real code, tested with injected clients; the release made no live provider calls, and the offline mode this page ports is the one that runs without any.
The source publishes no relevance or recall figures, only its fixture outcome, which this page reproduces: channel counts 3, 3, 5, 9, and note-a, note-p, note-e, note-c selected.
Limits in detail
Eighteen fictional notes and lexical features: no embeddings, no learned reranker, no LLM planner or judge, and no relevance judgments, so nothing here measures search quality. Tokens are ASCII and exact; there are no synonyms or stems. The corpus is the source's example and is not editable here; the query and settings are. An exported file is checked for consistency, not authenticity.
Reproduce it
Pick a row of the ladder, then open “Edit the query, see the scores, export a run”: drop the threshold to 3 and watch the shortfall; switch to body only; add a hard criterion that half the notes miss. Export a run and import it: the query reruns and the file is refused if its results do not follow. Files are limited to 50 KB. The port, the ladder and the tests are in the code for this page. From the site's Next.js app:
npx vitest run src/lib/projects/staged-search src/components/projects/staged-search