paniolo qmd Guide Research
Paper → Retriever → Product

The bibliography
behind retrieval

paniolo qmd searches a local index of skills, docs, and wiki pages, then hands back a path you can open. This page is why the stack is keyword first, vectors second, and a model only when one is already loaded. Commands stay on the qmd guide.

How To Read This Page

Their collections, not this index

None of these papers measured a Paniolo workspace. BM25, dense retrieval, fusion, HyDE, and cross-encoder reranking were studied on TREC, Wikipedia, MS MARCO, and the BEIR suites. qmd bench and qmd eval can score an index you own. This page does not publish a number from that. Where a paper's mechanism is optional or absent — a generator for query expansion, a reranker, an embedded index — the paragraph says what runs instead.

The Floor

A keyword search
that needs no model

BM25

BM25 scores a document by how often the query's terms appear in it, discounted by document length and by how common each term is in the collection. A rare symbol outranks a word that is everywhere. Robertson and Zaragoza's review is the account of that formula and of why it became the default lexical ranker. paniolo qmd search is that ranker over the local index: no embedding, no reranker, no network. A slug, a filename, or a function name is the query it is for.

BEIR then showed why a lexical floor is not a nostalgia feature. Across heterogeneous retrieval tasks, a neural ranker that looked strong on the collection it was trained on often failed to beat BM25 once the documents changed. In-domain gains did not transfer. Paniolo's hook takes that as an operating rule. The default prompt search is hybrid, and when the warm server or the vectors are unavailable it falls back to keyword search. A search that cannot run injects nothing, and the prompt still proceeds. We have not run BEIR on a customer index, and we do not claim BM25 wins every query — only that it is the path that still answers when no model is loaded.

Robertson & Zaragoza — Foundations and Trends in IR, 2009, doi.org/10.1561/1500000019 · Thakur et al. — BEIR, arXiv 2021, arxiv.org/abs/2104.08663
When The Words Differ

Vectors, fusion,
and a hypothetical page

Dense passage retrieval

Dense passage retrieval embeds the query and the documents in the same vector space and returns the nearest passages. The point of the embedding is vocabulary mismatch: the page says "session cookie" and the question says "how login persists." Karpukhin and colleagues showed that a dual encoder, trained on question–passage pairs, could retrieve those passages by similarity rather than by shared tokens. paniolo qmd vsearch is that arm alone. It needs an index you built with qmd embed. Without vectors, the command says so instead of pretending a keyword search was semantic.

Karpukhin et al. — Dense Passage Retrieval, arXiv 2020, arxiv.org/abs/2004.04906

Reciprocal rank fusion

Keyword rank and vector rank disagree, and their raw scores are not on the same scale. Reciprocal rank fusion ignores the scores and combines the ranks: a document's fused score is the sum, over each list, of one over (k + rank). Cormack, Clarke, and Buettcher found that this simple combination beat both the individual lists and a more elaborate voting method, and the constant they used was k = 60. paniolo qmd query fuses the BM25 list and the vector list with that constant. The arms that searched the words you typed are weighted above any model-written rewrite, so a handful of plausible expansions cannot outvote the question that was actually asked. query --explain prints where each arm placed a hit. Their SIGIR result is on TREC-style lists. It is not a measurement of this fusion on your wiki.

k = 60 — Cormack, Clarke & Buettcher, SIGIR 2009, doi.org/10.1145/1571941.1572114

HyDE, and query expansion

HyDE asks a model to write a short passage that would answer the question, then embeds that passage instead of the question. A hypothetical answer sits closer to the documents than the question does, which is why a zero-shot dense retriever can work without relevance labels. Gao and colleagues measured that on retrieval benchmarks. In qmd the same idea is a field you can write yourself — hyde: on a structured query, a 50–100 word passage in the voice of the page you hope to find — and a variant the expander may emit when a generator is configured. Expansion also proposes extra lexical and semantic wordings. It is skipped when no generator is configured, when the deadline cannot afford it, or when the call fails. The query then runs on the original arms. Expansion is a widening of the search, not a license to rewrite the wiki. That refusal is the same one on the wiki research page.

Gao et al. — HyDE, arXiv 2022, arxiv.org/abs/2212.10496
The Last Pass

Rerank a small pool,
or don't

Cross-encoder reranking

A bi-encoder scores the query and the document separately, which is cheap enough to run over the index. A cross-encoder reads them together, which is more accurate and too expensive to run over every page. Nogueira and Cho's passage re-ranker is that second stage: retrieve a candidate pool with a fast method, then rescore the pool with a model that sees the query and the text at once. query does this only when a reranker is configured. The pool is a capped set of fused candidates, and each candidate is itself capped at a few chunks, because every chunk is a full forward pass. Candidates the budget never reaches keep their fusion order and carry no rerank score. --no-rerank skips the stage. vsearch never enters it. The paper's gains are on its own re-ranking sets. A CPU-only machine that skips the cross-encoder has not failed the design. It has taken the cheap path the paper assumes exists.

Nogueira & Cho — Passage Re-ranking with BERT, arXiv 2019, arxiv.org/abs/1901.04085
What The Agent Receives

A path and a passage,
not the corpus

Retrieval, then the document

Lewis and colleagues' retrieval-augmented generation retrieves passages and conditions a generator on them, so the model can answer from a store it was not trained on. The useful half for a coding agent is the retrieval. The half Paniolo keeps narrow is what happens next. search, vsearch, and query return a path, a score, and a snippet. qmd get prints the document, line-numbered, so the agent quotes a span from the file rather than from a generated paraphrase. The prompt hook is the augmented-context use: a task-shaped prompt is searched, and the hits are injected before the agent answers. Short prompts and status questions inject nothing. That is a few hits in front of one prompt. It is not a generator standing in for the page.

Lost in the Middle is why the injection stays small. Liu and colleagues found that models use evidence at the start and end of a long context and miss what sits in the middle, even when the window is large enough to hold it. Stuffing the wiki into the prompt spends the budget the wiki research page traces to Codified Context, and it hides the page that mattered. The compiled store stays on disk. Search returns a pointer. The agent opens the file.

Lewis et al. — Retrieval-Augmented Generation, arXiv 2020, arxiv.org/abs/2005.11401 · Liu et al. — Lost in the Middle, arXiv 2023, arxiv.org/abs/2307.03172