The loop
behind evolve
The paper's numbers stay on harness research. This page is the part of that bibliography the CLI actually runs: inventory the harness, propose an edit, keep it only if validation holds, and roll it back when it does not.
The shape, not the campaign
Lin and colleagues ran ten unattended iterations and published a pass@1 lift. Paniolo has not reproduced that campaign, and this page does not repeat the percentages. What follows is the smaller loop paniolo evolve ships, and the places it stops short of the paper: no layered debugger over raw traces, no default that writes a diff you did not ask for.
See it, try it,
keep it or drop it
Diagnose
Component observability, in the paper, means every editable harness piece is a file, so a failure maps to one place and a change is a diff. paniolo evolve diagnose is the read-only start of that idea. It inventories the components it can see, reads trajectory logs, and emits a gap report. The report is a nomination. It is not the paper's debugger corpus, which distills millions of raw trajectory tokens into a layered evidence store an editor agent consumes. A gap report you can read is the piece that ships. The distillation pipeline does not.
Run
Decision observability pairs an edit with a prediction and checks it on the next round. paniolo evolve run walks the local version of that round: diagnose, a hypothesis, validation, then commit or rollback. The default proposer is offline, so a run does not call out for a model unless you choose one. --dry-run prints the diff and applies nothing. One iteration is the default, not ten. None of that is an unattended reproduction of their Terminal-Bench campaign. It is the same sequence, stopped where a person can still see the diff.
Judge
The paper's own limit is that the loop attributes fixes more reliably than regressions, and that stacking edits which each looked good does not add up. Lin and colleagues measured it. Fix predictions landed at 33.7% precision and 51.4% recall, about five times their random baseline. Regression predictions landed at 11.8% precision and 11.1% recall, about twice random. Across nine rounds the evolve agent named 43 regressions, five of them landed, and 40 actual regressions were not foreseen. Those figures are the paper's, on their campaign.
JUDGE is the CLI's answer to that limit. After validation, it compares the suite with the pre-edit baseline and rejects an edit that drops the pass rate, fails tasks that passed before, or blows the token budget past the configured policy. A rejected edit is rolled back. Acceptance does not rest on the proposer's own safety story. The check is a gate on the suites you registered. It is not the paper's per-edit prediction, verified task by task on the next benchmark round, and it does not claim to see the regressions their self-attribution missed.
Rollback
Anthropic's long-running harness treats git as the memory that survives a context window: leave a commit the next session can resume, and return to a known tree when the increment was wrong. paniolo evolve rollback is that return, applied to the harness component set rather than to the whole repository. It restores a prior snapshot, last-stable unless you name another decision. The initializer agent and the one-feature-per-session coding agent from that post stay on the harness page. They are not commands. The command is the snapshot.
Other loops,
and what we refuse
Meta-Harness
Meta-Harness searches harness code with a coding agent as the proposer and the filesystem as the memory. Every candidate is a directory of source, scores, and full traces. The proposer greps that history instead of receiving a crushed summary — on the order of 10 million tokens available per iteration, against the few thousand tokens earlier prompt optimizers actually consumed. On Terminal-Bench 2 the paper reports Opus 4.6 at 76.4% and Haiku 4.5 at 37.6% for the harness that search found. Those are their runs.
The piece paniolo-evolve takes is the navigable store: a hypothesis should be able to open the trace, not only a score. The piece it does not take is the unattended population search. The default proposer is offline. A coding-agent proposer is opt-in, and --dry-run still prints the diff. Their leaderboard numbers are not a result of this CLI.
Self-Harness
Self-Harness keeps the proposer and the task agent as the same frozen model. It mines failed traces into named failure patterns, proposes a small set of minimal edits tied to those patterns, and promotes an edit only after regression tests. Held-out pass rates on Terminal-Bench 2.0, from a minimal starting harness, moved MiniMax M2.5 from 40.5% to 61.9%, Qwen3.5-35B-A3B from 23.8% to 38.1%, and GLM-5 from 42.9% to 57.1%. The edits were model-specific. Lengthening the prompt was not the pattern.
That promotion rule is the same shape as JUDGE: non-regressive or it does not land. The mining loop is not the default. paniolo-evolve does not assume the model under test should be the one rewriting its own harness, and it does not treat those held-out lifts as something this product has measured.
Rethinking the evaluation
Wang and colleagues argue that a harness-evolution paper can report a gain that is really extra search. Evolution spends a budget on repeated evaluate-and-revise steps. Parallel sampling and sequential refinement spend a budget too. Unless those are matched, a higher pass@1 does not say the harness got better. A second failure is using the same public benchmark to search and to report. On Terminal-Bench 2.1, with Claude Opus 4.6, GPT-5.4, and GPT-5.4 mini, automatic harness evolution did not consistently beat simple test-time scaling — without unit tests, with unit tests, and when the evaluation tasks were not the search tasks.
That is why this page does not repeat the 69.7% to 77.0% figure as an evolve result. The figure lives on harness research, labeled as Lin and colleagues' campaign. A paniolo evolve run that improves a suite it just searched has not cleared the bar this paper sets. Promotion belongs on a held-out view, and a lift still has to be compared with spending the same budget on more samples.
VeRO
VeRO is an evaluation harness for an agent that optimizes another agent. Every edit is a snapshot with a diff and a rollback. A budget blocks further evaluations once it is spent. Traces, scores, and history arrive as a standard observation, not as a paragraph the optimizer wrote about itself. Held-out data and the model configuration stay outside the optimizer's write scope.
That contract is the run store around JUDGE. Validation spend is a gate, not a sentence in the prompt. The part VeRO does not settle, and this CLI does not pretend to, is which statistical test a small suite can support. A handful of tasks cannot carry a precise pass-rate claim. The gate can still refuse a regression it can see.
Where the edit is allowed to land
MOSS splits self-evolution into text-mutable artifacts — skills, prompts, memory, workflow graphs — and the source harness: routing, hooks, dispatch. A prompt edit cannot repair a hook-order bug, and a source edit takes effect whether or not the model feels like complying. Their pipeline anchors each change to a batch of production failures, lets an external coding agent write the diff, replays the batch, and promotes only with a person's consent. ReCreate makes the complementary point about evidence: a scalar score does not say why a scaffold failed, so the optimizer should open the trace, and a fix for one case should not become the domain agent until it generalizes.
paniolo-evolve keeps both cautions and ships neither system. Component edits are files you can diff. The proposer does not rewrite the product's own routing in the background, and it does not mint a new specialist agent from a trivial seed. A one-task patch stays a hypothesis until validation, on a suite that is not the conversation that suggested it, says otherwise.
MOSS — arxiv.org/abs/2605.22794 · ReCreate — arxiv.org/abs/2601.11100