The bibliography
behind the harness
Agentic Harness Engineering is the paper. OpenAI and Anthropic wrote the practitioner accounts beside it. This page says what each one argues. The short summary stays on the homepage. How paniolo evolve runs the loop is on evolve research.
Their experiment, not ours
The percentages below are Lin and colleagues' results, on their benchmarks, with their seed harness and their evolution loop. Paniolo has not reproduced them. The OpenAI and Anthropic posts are engineering accounts of how those teams build, not effect sizes. The CLI loop — diagnose, judge, commit or roll back — is on evolve research, including the campaign and the trace debugger this product does not ship.
Hold the model fixed,
edit the harness
Agentic Harness Engineering — Lin et al.
AHE treats the harness as the thing you can edit while the base model stays frozen: system prompt, tool descriptions and implementations, middleware, skills, sub-agent configuration, and long-term memory. Each of those is a file, so a change is a diff you can revert. The loop has three observability pillars. Component observability makes that file set the action space. Experience observability distills raw trajectories into a layered evidence corpus an editor can read, instead of asking it to swallow millions of tokens of logs. Decision observability pairs every edit with a prediction, then checks that prediction against the next round's task outcomes. An edit that fails the check is reverted. The paper's own limit is worth keeping: the loop attributes fixes more reliably than regressions, and stacking edits that each looked good does not add up.
On Terminal-Bench 2, ten iterations raised pass@1 from 69.7% to 77.0%, past the human-designed Codex harness at 71.9% and past the self-evolving baselines they compare (ACE and Training-Free GRPO). The frozen harness, without another evolution pass, transferred to SWE-bench-verified at the paper's highest aggregate success and 12% fewer tokens than the seed. On Terminal-Bench 2 it also gained +5.1 to +10.1 percentage points across three other model families. An ablation put the gain in tools, middleware, and long-term memory. The system prompt alone regressed. The quote on the homepage — tools, middleware, and memory carry the gain — is that ablation, not a Paniolo measurement.
paniolo scan reads the harness as files and scores it. It never writes, which is the diagnostic half of "if the harness moves outcomes, the harness deserves a score." The trace from scan's rules back to other papers is on scan research. The evolution loop that would edit those files — diagnose, a hypothesis, validation, commit or rollback — is paniolo evolve.
The ablation is why the rest of the product is split the way it is. Always-loaded prose stays short, and paniolo scan budgets it, because the paper found the system prompt was the weak component. Skills are a catalog loaded by name. The wiki is the long-term memory: markdown in the repo, linted, and searched with paniolo qmd rather than pasted into the prompt. Those are the layers the ablation said were carrying the result. They are not a claim that a Paniolo harness has been shown to move pass@1.
What the labs wrote
beside the paper
Harness engineering — OpenAI
OpenAI's account of building with Codex is a constraint story, not a benchmark table. The team stopped treating code-writing as the job and started treating the environment as the job: intent specified in the repo, feedback loops the agent can run, and knowledge laid out where an agent can find it. Design docs, execution plans, and a quality grade over domains live in a docs tree. AGENTS.md is a map, not the knowledge base. Custom linters enforce structural rules, and the lint message itself tells the agent how to fix the finding. When human review became the bottleneck, they made the UI, the logs, and the metrics legible to the agent instead of asking people to read more output.
That is the same split Paniolo ships, without their repository. paniolo scan is the linter over the harness, and the finding is the remediation hint. The wiki holds the design notes and decisions so they are not stuffed into AGENTS.md. paniolo stale is the feedback loop on those pages: a code change nominates a claim that may have gone false, and a person or an agent adjudicates it. The post does not measure a scanner. It is why the scanner's job is to make the environment legible, and why a finding that cannot name a fix is an incomplete finding.
Effective harnesses for long-running agents — Anthropic
Anthropic's harness for work that outlasts one context window is two roles and a clean handoff. An initializer session lays down the environment: a startup script, a progress file, and an initial git commit. Later sessions do one increment, then leave the tree in a state the next session can resume — a commit with a real message, and a progress note. The failure they are designing against is an agent that tries to finish the whole task in one window, or that leaves a mess the next window cannot see. Git is the memory that survives compaction.
Paniolo's counterpart is the files between sessions, not a pair of agent prompts. The wiki log and the page history are the progress note. paniolo stale records which claims a commit put in doubt, so the next session inherits an allegation instead of a silent drift. Returning the harness component set to a known snapshot is evolve rollback. We do not ship their initializer agent or their one-feature-per-session coding agent. The lesson we take is narrower: a session that cannot leave a file and a commit has not finished, even if the model thinks it has.