paniolo qmd Guide Research
CLI guide

paniolo qmd

Search the skills, docs, and wiki pages in a harness locally, then open the hit. Keyword search needs no model. Vector and hybrid search use an index you build on your machine.

Quick start

Index once, then search

From the harness root — the directory that holds paniolo.config.json — build the index, then look something up. The CLI walks up from the working directory to find that root. --root or PANIOLO_ROOT names it explicitly.

npx @paniolo/cli qmd reindex npx @paniolo/cli qmd search "vitest hook mock" npx @paniolo/cli qmd get "#abc123"

reindex refreshes the text index and, when this machine can embed, the vectors. search prints one block per hit — path, title, score, and a snippet. The result listing shows every line. get prints the document, line-numbered, so you can quote a span.

A missing flag on an older install means that binary predates it. paniolo qmd --help is the list for the copy you are running.

Search modes

Keyword, vector, or both

Three commands share one result shape. Pick by what you already know.

CommandUsesReach for it when
searchBM25 keywords. No model.You know the term, slug, filename, or symbol.
vsearchEmbeddings only. Alias vector-search.You can describe the idea and the words in the doc may differ. Needs an embedded index.
queryKeywords and vectors fused, then reranked when a reranker is configured. Alias deep-search.You want the best ordering and can pay for the model load.
paniolo qmd search "zustand store hook" paniolo qmd vsearch "how the warm sidecar avoids a second spawn" paniolo qmd query "effect-ts error handling in scripts"

vsearch and query say so when the index has no vectors — run qmd embed or qmd reindex first. query --no-rerank ranks on the fusion alone and skips the cross-encoder, which is the costly stage on a CPU-only machine. -C caps how many chunk passes the reranker may score.

query --explain prints where the keyword arm and the vector arm each placed a hit, what they contributed to the fused score, and the rerank score when one was computed. --explain, --no-rerank, and -C run in this process: the warm server serves one fixed pipeline and cannot vary it per call. The command says so on stderr when that happens.

Query craft

Name the intent, the words, and the idea

A bare sentence works. A structured query works better, because you choose which signal is lexical and which is semantic. Put the most important line first — it is weighted twice. Write intent: plus at least one of lex:, vec:, or hyde:.

FieldWhat to put there
intent:What you are trying to find, including the nearby topic to steer away from.
lex:Exact terms, skill slugs, filenames, symbols.
vec:The idea in the kind of language the source itself uses.
hyde:A short hypothetical passage that would answer the question, about 50–100 words.
paniolo qmd query "intent: harness vitest setup, away from playwright e2e lex: vitest mock hook test best-practices vec: how to unit test a react hook with vitest hyde: The vitest skill explains mocking a hook and asserting state transitions with renderHook."

--intent "…" prepends an intent: line to whatever query you passed. A single rare token or a verbatim phrase belongs on search.

Results

What a hit looks like

search, vsearch, and query print the same text block. Hits are separated by a blank line. This is one hit:

qmd://my-app/wiki/decision-auth.md:40 #abc123 Title: Auth decision Context: Markdown across the my-app repo. Score: 86% The session cookie is issued only after the IdP callback…
LineMeaning
qmd://…:40 #abc123Collection, repo-relative path, and the line where the match sits. #abc123 is the docid get accepts.
TitleThe document title, when the file has one.
ContextWhich collection the hit came from, in a short phrase such as "Markdown across the my-app repo."
ScoreRelevance as a percent of a 0–1 score. 86% is 0.86. --min-score uses that 0–1 scale.
SnippetThe passage around the match. --full replaces it with the document body.

Three further lines appear only when they apply:

LineWhen it appears
Status: completedThe page declares a lifecycle status: in frontmatter.
Archived: wiki page snapshotThe hit is a frozen copy under raw/wiki-archive/.
Duplicates: also indexed at 2 other path(s)The same bytes are indexed in more than one collection.
Calibration: below the relevance floor — a sampled probe, may not be relevantThe hit scored under the relevance floor and was included as a probe.

--format json returns the same fields under results, with score still on the 0–1 scale:

{ "results": [ { "docid": "abc123", "file": "qmd://my-app/wiki/decision-auth.md", "title": "Auth decision", "context": "Markdown across the my-app repo.", "line": 40, "score": 0.86, "snippet": "The session cookie is issued only after the IdP callback…" } ] }

--format files prints one path per line. The other formats and the flags that change this block are in the flag table.

Result shape

The flags all three modes share

FlagWhat it does
-n, --limitMax hits. Default 5. Default 20 for --format json and --format files.
--allEvery match, capped so a broad query cannot load the whole corpus.
--min-scoreDrop hits below this, on the 0–1 scale shown as a percent.
-c, --collectionLimit to these collections. Repeat the flag; the lists merge by score.
--fullPrint each hit's body. --line-numbers numbers those bodies.
--full-pathFilesystem paths in place of qmd:// URIs.
--historyKeep raw score order. Closed and archived wiki pages otherwise rank last.
--formatcli (default), json, csv, md, xml, or files (one path per line). --json, --csv, --md, --xml, and --files are the same choices.
-v, --verboseGlobal. Print model-load diagnostics on stderr.

Unquoted words are joined, so qmd search zustand store is the same query as qmd search "zustand store".

Scope

Collections follow the workspace

An unscoped search covers every collection generated from the harness workspace. Each configured repo is a collection named for that repo. A hidden subtree such as .agents/skills is its own collection, <repo>-agents-skills, and is left out of the repo collection. Scoping with -c my-app therefore misses that repo's skills; add -c my-app-agents-skills, or search unscoped, when you want both.

paniolo qmd ls paniolo qmd ls my-app paniolo qmd search "auth flow" -c my-app -c my-app-agents-skills

ls with no argument lists collections. ls <collection> or ls <collection>/<prefix> lists indexed documents. doctor prints each collection with its document count when you are unsure of a -c name. A name that is not indexed matches nothing, and the command names the collections that are.

Read the hit

Fetch by docid, path, or line window

Search snippets are a map. get and multi-get are the page. Output is line-numbered; --no-line-numbers turns that off. A :from:count suffix, or --from and -l, returns a window. --from is 1-indexed and overrides the suffix.

paniolo qmd get "#abc123" paniolo qmd get "#abc123:120:40" paniolo qmd get "qmd://my-app/wiki/decision-auth.md" --full-path paniolo qmd multi-get "#abc123,#def456"

multi-get also accepts a glob or a comma-separated list of paths. --max-bytes skips a document larger than the cap. --json emits the same retrieval as JSON.

Index

Rebuild after the corpus changes

The index is generated from the workspace file named in paniolo.config.json. After you add or edit guidance, refresh it. Windows and WSL each keep an index on their own native disk, so the two never share one file.

CommandWhat it does
updateRefresh the text index.
embedGenerate vectors. Uses the GPU when this machine has one.
reindexupdate, then embed when embedding is enabled. --prune drops collections that left the workspace. --background schedules the work and returns.
cleanupRelease cache and orphaned rows. Cheap. --dry-run reports without deleting.
vacuumCompact the index file and return free pages to the disk. Takes an exclusive lock. --dry-run reports the reclaimable size.

cleanup leaves the file size alone. vacuum is the rewrite that shrinks it.

Runtime

Status, the GPU, and the warm server

status reports what is indexed and whether a warm server is up. doctor diagnoses the stack: config, index, vectors, the warm server, and the device. Add --json to either. doctor --models loads the embedding model and reports how much of it sits on the GPU — that load is why the flag is off by default.

paniolo qmd status paniolo qmd doctor paniolo qmd gpu status paniolo qmd gpu 8 paniolo qmd serve --ensure

gpu reads or writes the per-machine preference in .qmd-local.json. on offloads every layer, off none, full also places the token embedding matrix on the device, and auto defers to the platform. A number offloads that many layers — 8 is the useful setting on an integrated GPU, where offloaded tensors still stage through system RAM. status reports without writing.

A ready warm server answers query in seconds. serve --ensure starts one when none is healthy and then exits. serve --restart replaces it. serve --stop asks it to shut down. serve --stop --all also reclaims servers that name no harness. Stop the server this way: a process-wide kill also takes down the MCP server.

ps lists the model daemons and sidecars on this machine (--json for the inventory). daemon --stop stops the only running pool, or daemon --stop <pool> names one. Model execution lives in paniolo-inference-worker, a pinned artifact the CLI installs and reuses. The CLI does not compile it.

Hooks

The harness searches before the agent does

A stamped harness registers paniolo qmd hook with each coding agent. The agent runs that command when a prompt, session, or edit event fires, and qmd writes context back on stdout. paniolo init includes these hooks. --no-retrieval on an adopt leaves only the safety hooks. On a harness that already exists:

paniolo evolve hooks install --retrieval --vendor claude paniolo evolve hooks install --retrieval --all

--dry-run prints the plan first. A hook you added yourself is left in place. The command the vendor actually runs looks like this, with the event payload on stdin:

paniolo qmd hook --vendor cursor --event user-prompt

Vendors are claude, cursor, codex, copilot, devin, and antigravity.

Each prompt

--event user-prompt runs on the prompt you just sent. A task-shaped prompt is searched and the hits are injected as additional context, so the agent sees relevant skills and docs before it answers. The default mode is hybrid query, with keyword search as the fallback when the warm server or the vectors are unavailable. Set qmd.hooks.mode to search for keywords only, or QMD_HOOK_MODE for one shell. The same context is mirrored into .cursor/rules/qmd-retrieval.mdc and removed again when a prompt produces none.

Short prompts and status questions ("did the tests pass") skip the search and inject nothing. A search that cannot run also injects nothing, and the prompt still proceeds. Antigravity receives the augmented prompt as plain text; the other vendors receive hook JSON.

Session start and maintenance

--event session-start injects guidance: this workspace has qmd, search it before task work, and a separate prompt hook does the compact per-prompt search. It does not query the index itself. Replace that paragraph with qmd.hooks.sessionStartGuidance in paniolo.config.json.

--event session-maintenance starts a background reindex and makes sure a warm server is up, then returns immediately so the session is not held for the model load.

After an edit

--event post-edit refreshes the index when a markdown file changes. A warm server folds that into its own watcher. Otherwise the hook starts one coalesced reindex and returns. A refresh that fails is reported on stderr and does not fail the edit. Reads do not reindex. When staleness is explicitly enabled, the same hook can also surface open allegations on the files just edited.

Command reference

What you run, and what stays in the background

CommandRole
search, vsearch, queryFind pages.
get, multi-getRead a hit, or several.
ls, status, doctorSee collections, index health, and the runtime.
update, embed, reindexRefresh the index.
cleanup, vacuumReclaim space.
gpu, serve, ps, daemonAcceleration and the long-running helpers.
mcpAlias of paniolo mcp. Editors start it; you rarely type it.
hookWhat the harness hooks call. Prompt search, session guidance, and the post-edit refresh.
bench, eval, tuneScore retrieval against a fixture, logged hook queries, or later edits. For one query you just ran, use query --explain.
For coding agents

Give your agent the qmd skill

The paniolo-qmd skill tells an agent to search before it reads, to author lex: and vec: itself, and to fetch with get before it quotes a page. Install it with the rest of the catalog:

npx skills add paniolo-ai/skills --skill paniolo-qmd

The corpus qmd searches is the LLM Wiki plus the rest of the workspace guidance. MCP exposes the same retrieval as tools. Collection membership is the qmd field in the config reference.