The bibliography
behind the knowledge base
Compiled wikis, a walker trained to use one, and stores that change because a question was asked. The pattern itself — raw sources, linked pages, a short schema — stays on wiki research. Commands stay on the wiki guide.
Their benchmarks,
their stores
Every number below is from the paper that measured it. Paniolo has not run these suites. The product that exists is a directory of markdown pages, a linter, a log, and paniolo qmd as the read path. An Error Book, a trained navigator, and a store that rewrites itself from gold answers are described here because the papers measured them. They are not commands.
Pages an agent
can walk
LLM-Wiki — retrieval as reasoning
Ming, Li, Wu, and Que compile a document corpus into linked markdown pages and give the agent three tools: search, read, and follow a link, until it decides the evidence is enough. An Error Book records a recurring construction mistake — a dangling link, an unsupported fact, a contradiction across pages — and turns it into a constraint for the next batch. Structural fixes are deterministic. Semantic fixes are a later model pass. Karpathy's gist named the pattern. This paper is a system with a score.
On the first 500 questions of HotpotQA, MuSiQue, and 2WikiMultiHopQA, over the context paragraphs that come with those questions, LLM-Wiki's answer F1 is 0.839, 0.739, and 0.911. The paper reports that as +2.0, +8.1, and +6.4 F1 over LightRAG, and the gap grows as the question needs more hops. That corpus is the dataset's own paragraphs, not open-domain Wikipedia, and the agent may spend up to 15 tool calls. On AuthTrace, judged by GPT-4o-mini, they lead HippoRAG 2 by 2.1 points overall and by more on multi-document questions (+5.0 and +8.9). On single-document questions HippoRAG 2 leads by 2.3. A page that summarizes a source can drop the sentence a one-hop question needed.
paniolo wiki compiles pages, checks links, and keeps log.md. It does not keep an Error Book, and qmd does not decide that the evidence is sufficient. The single-document regression is the caution: a compiled page is a better place to synthesize, and a worse place to hide a local detail.
ByteRover — the agent writes the files
ByteRover inverts the usual memory pipeline. The model doing the task also writes the memory, as a tree of markdown files: domain, topic, subtopic, entry. Each entry carries relations, provenance, and a lifecycle from draft to validated to core. Retrieval is a local full-text index that usually answers without another model call. There is no vector database and no graph database. Writes are tools — add, update, merge, delete — and the agent sees what each one did.
On LoCoMo, 1,982 questions, their harness, Gemini 3 Flash as judge, overall accuracy is 96.1 against HonCho at 89.9 and Hindsight at 89.6. Multi-hop is 93.3 against HonCho's 84.0. Open-domain is the miss: Hindsight 95.1, ByteRover 85.9. Those questions need something the conversation never said. On LongMemEval-S their overall is 92.8, with multi-session the weak category at 84.2. Chronos-High, at 95.6 with Claude Opus 4.6, is a different generator. Rows in their table that use another paper's judge are not the same comparison.
The file shape is the part Paniolo already has: markdown, a tree, provenance, operations that report what they changed. qmd will embed when it is asked to. ByteRover's claim is that it never does. The 96.1 is a conversational-memory score under their judge. It is not a lint score.
ByteRover — arXiv 2026, arxiv.org/abs/2604.01599Who walks
the pages
SearchWiki — a trained walker
SearchWiki compiles a corpus into three layers — document overviews, cross-document topic pages, and page-level source records — and then trains WikiResearcher-9B, a Qwen3.5-9B policy, to descend that tree. Every model in the main table uses the same tools. An untrained 27B or 397B in the table is a larger model walking the same wiki, not a flat chunk retriever.
On ViDoRe v3, strict accuracy macro-averaged over eight domains, WikiResearcher-9B scores 71.35 ± 0.73. The untrained 9B in the same harness scores 63.95 ± 0.76. The untrained 27B scores 70.94 ± 0.72 and the untrained 397B scores 69.51 ± 0.73. The paper's paired comparison puts the same-size gain at about +7.4 points. That is navigation skill on a fixed wiki. On FinanceBench, same store, the 9B policy scores 83.33, behind the untrained 27B at 85.33 and ahead of the untrained 9B at 79.33. The 2023 leaderboard figures they cite, all at or below 19.30, are a different setup. On LoCoMo token F1 the untrained 27B still leads, 55.84 to 48.31. A policy trained on document wikis can walk a dialogue wiki. It does not own conversational memory.
Paniolo's read path is qmd: keyword, vector when a model is loaded, and a short list injected into the prompt. There is no 9B policy trained to descend overviews, then topics, then sources. The layer split is still the right filing idea. An overview, a topic page, and the source record are different objects, and an agent that can stop at the overview spends less.
SearchWiki — Singh et al., arXiv 2026, arxiv.org/abs/2608.29953WikiLoop — the builder hears the navigator
WikiLoop trains the agent that builds the wiki and the agent that reads it as one policy. The builder proposes a structured edit. A frozen navigator scores that edit by the change in downstream utility, on the questions the edit was meant to help and on a guard set of unrelated questions. Regressions on the guard set are penalized. Accidental improvements there are not rewarded. The navigator's own reward waits until the evidence is complete before it charges for extra lookups.
On AuthTrace, with Qwen3.5-9B answering, WikiLoop's overall judged accuracy is 62.6 against an LLM-Wiki baseline of 56.3. The paper's splits versus that baseline are +2.4, +12.8, and +12.3 on single-document, low multi-document, and high multi-document questions. HippoRAG 2 still leads slightly on single-document questions, 69.8 to 69.1. The builder ablation is the curation result. Scoring a patch by how the wiki looks after the edit barely moves affected questions. Scoring the before-and-after difference raises that change from 4.9 to 12.8 and also raises guard regression to 0.046. The guard penalty cuts that regression to 0.016, a 65.2% relative drop, while the affected-question change only slips from 12.8 to 12.1. A held-out navigator, trained from a different seed and kept out of edit selection, still finds the guarded edits useful.
Paniolo does not train a builder from a navigator's score. The design that transfers without that run: an edit to the knowledge base is accountable to the questions it was meant to help and to the questions it must leave alone. A lint that only looks at the page you just touched will not see the second set.
WikiLoop — arXiv 2026, arxiv.org/abs/2607.26604Structure follows
the questions
Training a knowledge base
Pan and Yu treat the store as the model. A curator answers a supervised question from the current store, sees the gold answer, and then edits the store. At test time the store is frozen. A different reader, with no gold and a fixed action budget, answers held-out questions. GraphRAG, RAPTOR, and HippoRAG build structure from the corpus before any question arrives. Here the labels are question-answer pairs, and links are spent on questions that were actually asked.
Both arms are fictional universes, so the model cannot already know the facts. KBGym is their generator. PhantomWiki is someone else's. They did not run a real-text corpus. On KBGym the gain tracks how much of a test question the training set touched. Trained questions: +0.294 F1 over the flat store, and 25% fewer actions. Unseen questions whose both keys appeared in training: +0.176 F1, and the step saving is no longer significant. One shared key: +0.059 F1. No shared key: parity with the flat store. Their efficiency comparison with a HippoRAG-style graph, adapted into the same store and read by the same budgeted reader, is 1,913 links against 196,112, and per point of corpus covered, 1.5 times the action saving and 2.1 times the accuracy gain. That is not a rerun of HippoRAG's published QA tables.
A store that aces the questions it was edited for, and ties the flat store on questions that share nothing, has memorized a neighborhood. paniolo wiki does not train the corpus on gold answers. A search miss stays a retrieval result. The coverage gradient is the bar a future curation loop would have to clear before anyone called the corpus smarter.
WikiSkill — the wiki outlives the skill
WikiSkill keeps three layers. Raw traces are immutable. A wiki of pattern notes and an evolution log compiles what failed and what worked. Skills are the procedures the agent is allowed to keep. A validation split accepts a skill edit or rolls the skills back. The wiki stays, so the next proposal can see the rejection. During training rollouts the agent does not read the wiki. The authors' ablation says that access hurts skill development, and that dropping the persistent wiki hurts the result.
Across five tasks and three independent runs, the average gain over no skills grows with Qwen size: +12.3, +17.5, and +23.9 points at 4B, 9B, and 27B. Qwen-3.5-9B with the evolved skills averages 47.4, against Qwen-3.6-27B with no skills at 39.4. The 4B model with skills averages 38.5. Gemini on ALFWorld does not move, because the no-skill run already saturates the validation split. Those are skill-evolution scores. They are not a measurement of paniolo wiki.
The layer split is the one Paniolo already uses. raw/ is evidence. The wiki is compiled memory. Skills are a catalog loaded by name. Paniolo does not gate skill edits on a held-out validation split. The caution from their ablation is the one to keep: the agent that is practicing should not read the wiki it is supposed to be filling. The wiki is for the next edit.