paniolo qmd
Search the skills, docs, and wiki pages in a harness locally, then open the hit. Keyword search needs no model. Vector and hybrid search use an index you build on your machine.
Index once, then search
From the harness root — the directory that holds
paniolo.config.json — build the index,
then look something up. The CLI walks up from the working directory
to find that root. --root or
PANIOLO_ROOT names it explicitly.
npx @paniolo/cli qmd reindex
npx @paniolo/cli qmd search "vitest hook mock"
npx @paniolo/cli qmd get "#abc123"reindex refreshes the text index and,
when this machine can embed, the vectors.
search prints one block per hit — path,
title, score, and a snippet. The
result listing shows every line.
get prints the document, line-numbered,
so you can quote a span.
A missing flag on an older install means that binary predates it.
paniolo qmd --help is the list for the
copy you are running.
Keyword, vector, or both
Three commands share one result shape. Pick by what you already know.
| Command | Uses | Reach for it when |
|---|---|---|
search | BM25 keywords. No model. | You know the term, slug, filename, or symbol. |
vsearch | Embeddings only. Alias vector-search. | You can describe the idea and the words in the doc may differ. Needs an embedded index. |
query | Keywords and vectors fused, then reranked when a reranker is configured. Alias deep-search. | You want the best ordering and can pay for the model load. |
paniolo qmd search "zustand store hook"
paniolo qmd vsearch "how the warm sidecar avoids a second spawn"
paniolo qmd query "effect-ts error handling in scripts"vsearch and query
say so when the index has no vectors — run
qmd embed or
qmd reindex first.
query --no-rerank ranks on the fusion
alone and skips the cross-encoder, which is the costly stage on a
CPU-only machine. -C caps how many
chunk passes the reranker may score.
query --explain prints where the
keyword arm and the vector arm each placed a hit, what they contributed
to the fused score, and the rerank score when one was computed.
--explain,
--no-rerank, and
-C run in this process: the warm
server serves one fixed pipeline and cannot vary it per call. The
command says so on stderr when that happens.
Name the intent, the words, and the idea
A bare sentence works. A structured query works better, because you
choose which signal is lexical and which is semantic. Put the most
important line first — it is weighted twice. Write
intent: plus at least one of
lex:, vec:,
or hyde:.
| Field | What to put there |
|---|---|
intent: | What you are trying to find, including the nearby topic to steer away from. |
lex: | Exact terms, skill slugs, filenames, symbols. |
vec: | The idea in the kind of language the source itself uses. |
hyde: | A short hypothetical passage that would answer the question, about 50–100 words. |
paniolo qmd query "intent: harness vitest setup, away from playwright e2e
lex: vitest mock hook test best-practices
vec: how to unit test a react hook with vitest
hyde: The vitest skill explains mocking a hook and asserting state transitions with renderHook."--intent "…" prepends an
intent: line to whatever query you
passed. A single rare token or a verbatim phrase belongs on
search.
What a hit looks like
search, vsearch,
and query print the same text block.
Hits are separated by a blank line. This is one hit:
qmd://my-app/wiki/decision-auth.md:40 #abc123
Title: Auth decision
Context: Markdown across the my-app repo.
Score: 86%
The session cookie is issued only after the IdP callback…| Line | Meaning |
|---|---|
qmd://…:40 #abc123 | Collection, repo-relative path, and the line where the match sits. #abc123 is the docid get accepts. |
Title | The document title, when the file has one. |
Context | Which collection the hit came from, in a short phrase such as "Markdown across the my-app repo." |
Score | Relevance as a percent of a 0–1 score. 86% is 0.86. --min-score uses that 0–1 scale. |
| Snippet | The passage around the match. --full replaces it with the document body. |
Three further lines appear only when they apply:
| Line | When it appears |
|---|---|
Status: completed | The page declares a lifecycle status: in frontmatter. |
Archived: wiki page snapshot | The hit is a frozen copy under raw/wiki-archive/. |
Duplicates: also indexed at 2 other path(s) | The same bytes are indexed in more than one collection. |
Calibration: below the relevance floor — a sampled probe, may not be relevant | The hit scored under the relevance floor and was included as a probe. |
--format json returns the same fields
under results, with
score still on the 0–1 scale:
{
"results": [
{
"docid": "abc123",
"file": "qmd://my-app/wiki/decision-auth.md",
"title": "Auth decision",
"context": "Markdown across the my-app repo.",
"line": 40,
"score": 0.86,
"snippet": "The session cookie is issued only after the IdP callback…"
}
]
}--format files prints one path per line.
The other formats and the flags that change this block are in the
flag table.
The flags all three modes share
| Flag | What it does |
|---|---|
-n, --limit | Max hits. Default 5. Default 20 for --format json and --format files. |
--all | Every match, capped so a broad query cannot load the whole corpus. |
--min-score | Drop hits below this, on the 0–1 scale shown as a percent. |
-c, --collection | Limit to these collections. Repeat the flag; the lists merge by score. |
--full | Print each hit's body. --line-numbers numbers those bodies. |
--full-path | Filesystem paths in place of qmd:// URIs. |
--history | Keep raw score order. Closed and archived wiki pages otherwise rank last. |
--format | cli (default), json, csv, md, xml, or files (one path per line). --json, --csv, --md, --xml, and --files are the same choices. |
-v, --verbose | Global. Print model-load diagnostics on stderr. |
Unquoted words are joined, so
qmd search zustand store is the same
query as qmd search "zustand store".
Collections follow the workspace
An unscoped search covers every collection generated from the harness
workspace. Each configured repo is a collection named for that repo.
A hidden subtree such as .agents/skills
is its own collection, <repo>-agents-skills,
and is left out of the repo collection. Scoping with
-c my-app therefore misses that repo's
skills; add -c my-app-agents-skills, or
search unscoped, when you want both.
paniolo qmd ls
paniolo qmd ls my-app
paniolo qmd search "auth flow" -c my-app -c my-app-agents-skillsls with no argument lists collections.
ls <collection> or
ls <collection>/<prefix>
lists indexed documents. doctor prints
each collection with its document count when you are unsure of a
-c name. A name that is not indexed
matches nothing, and the command names the collections that are.
Fetch by docid, path, or line window
Search snippets are a map. get and
multi-get are the page. Output is
line-numbered; --no-line-numbers turns
that off. A :from:count suffix, or
--from and -l,
returns a window. --from is 1-indexed
and overrides the suffix.
paniolo qmd get "#abc123"
paniolo qmd get "#abc123:120:40"
paniolo qmd get "qmd://my-app/wiki/decision-auth.md" --full-path
paniolo qmd multi-get "#abc123,#def456"multi-get also accepts a glob or a
comma-separated list of paths. --max-bytes
skips a document larger than the cap.
--json emits the same retrieval as JSON.
Rebuild after the corpus changes
The index is generated from the workspace file named in
paniolo.config.json. After you add or
edit guidance, refresh it. Windows and WSL each keep an index on their
own native disk, so the two never share one file.
| Command | What it does |
|---|---|
update | Refresh the text index. |
embed | Generate vectors. Uses the GPU when this machine has one. |
reindex | update, then embed when embedding is enabled. --prune drops collections that left the workspace. --background schedules the work and returns. |
cleanup | Release cache and orphaned rows. Cheap. --dry-run reports without deleting. |
vacuum | Compact the index file and return free pages to the disk. Takes an exclusive lock. --dry-run reports the reclaimable size. |
cleanup leaves the file size alone.
vacuum is the rewrite that shrinks it.
Status, the GPU, and the warm server
status reports what is indexed and
whether a warm server is up. doctor
diagnoses the stack: config, index, vectors, the warm server, and the
device. Add --json to either.
doctor --models loads the embedding
model and reports how much of it sits on the GPU — that load is why
the flag is off by default.
paniolo qmd status
paniolo qmd doctor
paniolo qmd gpu status
paniolo qmd gpu 8
paniolo qmd serve --ensuregpu reads or writes the per-machine
preference in .qmd-local.json.
on offloads every layer,
off none,
full also places the token embedding
matrix on the device, and auto defers
to the platform. A number offloads that many layers —
8 is the useful setting on an
integrated GPU, where offloaded tensors still stage through system RAM.
status reports without writing.
A ready warm server answers query
in seconds. serve --ensure starts one
when none is healthy and then exits.
serve --restart replaces it.
serve --stop asks it to shut down.
serve --stop --all also reclaims
servers that name no harness. Stop the server this way: a process-wide
kill also takes down the MCP server.
ps lists the model daemons and
sidecars on this machine (--json for
the inventory). daemon --stop stops
the only running pool, or daemon --stop <pool>
names one. Model execution lives in
paniolo-inference-worker, a pinned
artifact the CLI installs and reuses. The CLI does not compile it.
The harness searches before the agent does
A stamped harness registers paniolo qmd hook
with each coding agent. The agent runs that command when a prompt,
session, or edit event fires, and qmd writes context back on stdout.
paniolo init includes these hooks.
--no-retrieval on an adopt leaves only
the safety hooks. On a harness that already exists:
paniolo evolve hooks install --retrieval --vendor claude
paniolo evolve hooks install --retrieval --all--dry-run prints the plan first.
A hook you added yourself is left in place. The command the vendor
actually runs looks like this, with the event payload on stdin:
paniolo qmd hook --vendor cursor --event user-promptVendors are claude,
cursor, codex,
copilot, devin,
and antigravity.
Each prompt
--event user-prompt runs on the
prompt you just sent. A task-shaped prompt is searched and the hits
are injected as additional context, so the agent sees relevant skills
and docs before it answers. The default mode is hybrid
query, with keyword search as the
fallback when the warm server or the vectors are unavailable.
Set qmd.hooks.mode to
search for keywords only, or
QMD_HOOK_MODE for one shell.
The same context is mirrored into
.cursor/rules/qmd-retrieval.mdc and
removed again when a prompt produces none.
Short prompts and status questions ("did the tests pass") skip the search and inject nothing. A search that cannot run also injects nothing, and the prompt still proceeds. Antigravity receives the augmented prompt as plain text; the other vendors receive hook JSON.
Session start and maintenance
--event session-start injects
guidance: this workspace has qmd, search it before task work, and a
separate prompt hook does the compact per-prompt search. It does not
query the index itself. Replace that paragraph with
qmd.hooks.sessionStartGuidance in
paniolo.config.json.
--event session-maintenance starts
a background reindex and makes sure a warm server is up, then returns
immediately so the session is not held for the model load.
After an edit
--event post-edit refreshes the
index when a markdown file changes. A warm server folds that into its
own watcher. Otherwise the hook starts one coalesced reindex and
returns. A refresh that fails is reported on stderr and does not fail
the edit. Reads do not reindex. When
staleness is explicitly enabled, the same
hook can also surface open allegations on the files just edited.
What you run, and what stays in the background
| Command | Role |
|---|---|
search, vsearch, query | Find pages. |
get, multi-get | Read a hit, or several. |
ls, status, doctor | See collections, index health, and the runtime. |
update, embed, reindex | Refresh the index. |
cleanup, vacuum | Reclaim space. |
gpu, serve, ps, daemon | Acceleration and the long-running helpers. |
mcp | Alias of paniolo mcp. Editors start it; you rarely type it. |
hook | What the harness hooks call. Prompt search, session guidance, and the post-edit refresh. |
bench, eval, tune | Score retrieval against a fixture, logged hook queries, or later edits. For one query you just ran, use query --explain. |
Give your agent the qmd skill
The paniolo-qmd skill
tells an agent to search before it reads, to author
lex: and vec:
itself, and to fetch with get before it
quotes a page. Install it with the rest of the
catalog:
The corpus qmd searches is the LLM Wiki plus
the rest of the workspace guidance.
MCP exposes the same retrieval as tools.
Collection membership is the qmd field
in the config reference.