The bibliography
behind the ledger
Every headline capability in paniolo stale traces to published freshness research or convergent practitioner diagnosis. This page is the trace — keyed by feature, not by paper — and the places we deliberately refuse the literature's shortcuts.
Evidence strength, labeled
The product docs cover commands and config. This page covers why the shape exists. Each claim carries one of four labels:
- Measured — a statistic or result from the paper's own study. The paper measured the problem or a method in its domain; it did not measure our detector's precision on your wiki.
- Taxonomic — the feature operationalizes a validated vocabulary or split from the literature (for example SUD vs SUNP, freshness vs age).
- Practitioner — engineering posts that independently converge on the same failure mode (docs and code lack a feedback loop; wrong prose is worse than missing prose). Motivation, not a peer-reviewed effect size.
- Paniolo ahead / behind — where the ledger has no close analogue in the ingested set, or where the literature offers a mechanism we have not shipped.
One claim we deliberately do not make: that retrieval scores, agent agreement, or a single model edit prove documentation truth. The central product rule — detection nominates; it never proves — is a response to negative results on automated knowledge editing, not a marketing flourish.
How candidates are nominated
Declared watches
Frontmatter staleness.watches ties prose to repository paths and symbols. A matching code diff nominates a falsifiable allegation — it does not close a verdict.
| Feeds | Anchors | Strength |
|---|---|---|
declared-watch/1, docs-declared/1 |
USB predict-then-refresh (IJCAI 2017); DEAN structural outdated-fact detection; Sync-o / Mintlify event-driven doc refresh | Taxonomic for predict-then-refresh; practitioner for change-triggered review. Watches are manual binary edges — not a learned update-frequency prior (behind the volatility literature). |
Exact deleted or renamed references
Lexical hits on deleted or renamed code paths and symbols can nominate unwatched prose. Exact identity only — not semantic drift.
| Feeds | Anchors | Strength |
|---|---|---|
unwatched-ref/1 |
DEAN outdated structured facts; Doccupine on trusted-but-wrong references | Taxonomic problem framing; detector is deterministic string identity |
Parser-proven code comments
Tree-sitter binds comments (and Python docstrings) to owning symbols. Ambiguous ownership creates no allegation. Remediation must preserve the executable token stream.
| Feeds | Anchors | Strength |
|---|---|---|
comment-assoc/1 |
SemUpdates text-level semantic update detection (SIGIR 2025); Sugarbug on structural disconnect between work and description | Taxonomic for text-level staleness. Comment↔code co-evolution papers exist outside this ingested set; association here is parser-proven engineering, not a Fluri-style mined link. |
Live-reference verification (the verifier role)
Agents judge claims against current code — or against an explicit claim-scope version — not against remembered snapshots.
| Feeds | Anchors | Strength |
|---|---|---|
| Verifier agent; claim-scope / historical markers | DyKnow live-reference checks; SemUpdates SUD/SUNP split; FreshLLMs evidence-at-answer-time; fact-duration / KB stability for time-bounded truth | Measured motivation from DyKnow's negative editing results; taxonomic for SUD vs SUNP and freshness/age vocabulary |
Durable work, not a one-shot detector
Falsifiable allegation queue
Content-addressed allegations, immutable evidence, reopen-on-edit, and retained insufficient-evidence. Unknown is not scored as fresh.
| Feeds | Anchors | Strength |
|---|---|---|
S- allegations, dispositions, scan checkpoints |
Cho & Garcia-Molina budgeted freshness; Ekline detect→draft→review→measure loop; Doccupine / Sync-o on partially-right pages being worse than missing ones | Taxonomic for freshness/age and budgeted revisit. The git-native allegation trail with optimistic revisions has no close analogue in the ingested paper set (ahead). |
Budgeted phases and deferred tail
maxAllegationsPerPhase, retrieval topK / maxImmediate, and durable deferred work are a refresh-budget problem in the Cho & Garcia-Molina sense.
| Feeds | Anchors | Strength |
|---|---|---|
| Worker budgets, deferred scheduling | Page refresh policies; streaming materiality / edit-velocity scores | Taxonomic budget framing. Deferred declared candidates are still allegations — scheduling is not shadow-only. |
Why four roles, not one model
Verifier + verdict challenger
Independent processes; the challenger sees the sealed verdict body, never the verifier's hidden reasoning. Disagreement retains work.
| Feeds | Anchors | Strength |
|---|---|---|
verifier, verdictChallenger |
DyKnow: automated verification/editing without opposition over-trusts the model; maker/checker harness practice | Measured caution from DyKnow. Adversarial separation of sealed verdicts is ahead of the ingested KB-staleness paper set (no close analogue). |
Remediator + patch challenger + merge gate
Remediation is a bounded rewrite (ERASE-style edit, not a stale note). Rust validates the unified diff in memory; merge requires upheld observations and an exact proposal head.
| Feeds | Anchors | Strength |
|---|---|---|
| Prose/comment remediation, merge gate, kill switch | ERASE rewrite-don't-append; WikiBigEdit limits of lifelong LLM editing; Mintlify/Ekline human review step (Paniolo substitutes dedicated challengers + repo gates) | Measured risk from WikiBigEdit / DyKnow. Allowlists and challengers bound blast radius — they do not claim editing is solved. |
Search nominates; it does not judge
qmd as a shadow lane
Lexical, vector, and rerank lanes may measure recall over a holdout. Scores never set dispositions, mint gold labels, or grant merge authority. Production nomination today is declared watches plus exact references until a refreshed holdout clears frozen gates.
| Feeds | Anchors | Strength |
|---|---|---|
shadow-qmd, live retrieval.shadow, index-integrity fail-closed |
FreshLLMs (fetch current evidence, don't rewrite weights); appropri8 freshness ≠ relevance; agent-ecosystem / Risingwave on index lag as pipeline staleness | Taxonomic for retrieval-as-refresh. Shadow remains non-authoritative by design — live shadow metrics are not admission. |
Demand and consumer signals (deferred)
Query-log and behavioral bounce signals appear in the literature and practitioner corpus. They are not a shipped scheduling prior — collection policy and measured benefit remain open.
| Feeds | Anchors | Strength |
|---|---|---|
Future demand prior; bounded flag observations |
Topic-aware KB update from query logs; Slite / Kipwise search-bounce; dbi embedding-version outcome metrics | Practitioner / taxonomic motivation. Behind as a production scheduler. |
Where the literature is ahead
Declared watches are a defensible deterministic floor. These are the largest research gaps relative to the ingested corpus — not active release promises:
- Learned volatility prior — fact duration, Wikidata stability (~83% F1 post-hoc in that study), materiality / Abnormal Edit Ratio, and per-entity update frequency all predict who goes stale before checking. The ledger records outcomes that could feed such a prior after opportunity-normalized calibration; it does not ship one yet.
- Typed propagation — WikiMonitor-Onto found one-hop ontology propagation surfacing indirect stale concepts a per-assertion baseline missed. Ordinary wikilinks do not currently fan out nominations.
- Remediator risk class — DyKnow and WikiBigEdit warn about LLM edits. Challengers and allowlists bound the blast radius; they do not eliminate the failure class.
Same diagnosis, shared loop
Independent engineering posts converge on the problem the ledger productizes: documentation and code live in disconnected systems; AI raises write velocity without review velocity; a wrong page is trusted where a missing page is not.
- Slite — knowledge drift
- Mintlify — documentation drift
- Ekline — detect, measure, prevent
- Sync-o — stale documentation
- Doccupine — keep docs up to date
- appropri8 — freshness-aware RAG
Those workflows usually place a person at approval. Paniolo places dedicated verifier, challenger, and remediator agents plus deterministic repository gates — still evidence-gated, still refuse-to-guess when ungrounded.
Read the operator guide, then try a dry-run scan.
npx @paniolo/cli stale scan — detection without agents