harness security Scan Secrets Research
Paper → Pattern → Product

The bibliography
behind harness security

Who is allowed to act, which setting is actually on, and what an unverified write can do later. Commands stay on the scan guide and the secrets guide. The scanner's own feature trace stays on scan research.

How To Read This Page

Their trials,
their harnesses

Every rate below is from the paper that measured it. Paniolo has not run these suites. paniolo scan reads a repository and scores what the harness declares. It never writes a file, and it does not sit on the tool-call path. A finding that an allowlist is missing is a declaration check. A runtime refusal, a kernel sandbox, and a trained guardrail are different products. This page names the mechanism and the published impact. It does not reproduce the trials.

Authority

The model cannot
mint permission

CapScope — authority is not a string

CapScope treats a coding agent's tools as ambient authority: naming a path or a command is enough to act on it. Untrusted text — a repository file, a skill, command output — can ask for a call the user did not request. The defense is a check the model cannot talk its way past. Before repository contents or tool output are read, a trusted preflight freezes a task-wide ceiling from the user's request and the project's file names. Each agent then holds its own capabilities, stored outside the context. A call runs only if the agent that issued it already holds a covering capability. Delegation can only narrow that set.

On Pi 0.78.0, five small Python repairs, five surfaces, four conditions, three trials: 300 runs. The injected effect is an ordinary in-repository action the repair does not need. It executed in 47 of 75 runs under ambient authority, 46 of 75 under a static denylist shared by every agent, and 33 of 75 under one task-specific allowlist still shared by every agent. CapScope admitted it in 3 of 75. All three remaining executions arrived through project instruction files. Repairs completed 68 of 75 under CapScope and 72 of 75 under ambient authority. The denylist fails because the actions are ordinary project operations the list never names. A shared allowlist can drop operations the task does not need, and still cannot give a write to the agent that patches a file while withholding it from the agent that only runs tests. The paper's own limit: five small tasks, one model, three trials. A scope minted too wide still admits an in-scope action.

The scan analogue is a declared tool allowlist, an untrusted-input boundary, and a confirmation step for a high-impact action. Those checks verify the protocol is written down. They do not refuse a call at dispatch, and the 3/75 figure is CapScope's, on their tasks.

CapScope — arXiv 2026, arxiv.org/abs/2609.08371

Open Agent Passport — a signed decision before the call

Open Agent Passport puts a policy decision on the before_tool_call hook: a declarative policy, and a signed audit of what was allowed. In the APort Vault CTF, a permissive policy let attackers succeed 74.6% of the time. A restrictive policy held at 0% across 879 attempts. The cloud API's median decision latency in their measurement was 53ms. The result is the case for enforcement over a paragraph that asks the model to be careful. Paniolo does not issue that passport. The hook shape is the part a harness can declare, and scan can see whether a high-impact action asks for confirmation.

Open Agent Passport — arXiv 2026, arxiv.org/abs/2603.20953
Settings

A setting that exists
and a setting that is on

HarnessSecurity-Bench — utility beside attack success

HarnessSecurity-Bench inventories ten mechanisms across 40 harnesses, 27 open and 13 closed: auto-approve, network isolation, command allowlisting, command denylisting, MCP permissions, prompt-injection filtering, path restriction, read-only mode, audit logging, and project trust. Of 400 cells, 205 are confirmed implemented, 83 are absent, and 112 are unresolved. About half of the confirmed implementations are opt-in. Closed-source harnesses have more cells the raters could not verify. Audit logging is in the taxonomy and out of the runtime bench, because a log does not by itself stop an effect.

The bench runs 23 tasks on six harnesses — Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, and GitHub Copilot — under GLM-5.2: 2,500 trials and 81,155 tool calls. Separate oracles score finishing the user's task and whether the extra effect landed. Turning auto-approve on raised attack success from 29.2% to 95.6%, and raised task utility with it (77.1% to 95.4% overall). Network isolation cut attack success from 57.1% to 0.8% and cost 24.5 percentage points of utility; almost all of that loss was work that needed the network. Read-only mode cut attack success from 8.3% to 0.8% and cost 34.0 points of utility. Command denylisting, path restriction, and MCP permissions cut attack success by 40.4, 48.1, and 28.1 points, with overall utility changes of +5.2, +0.3, and +0.5 points. Allowlisting is reported as a small utility loss. A restriction on a capability the legitimate task also needs will block both. A command the policy allows can still reach the protected operation by another path.

Scan already flags blanket auto-approve, hooks that call non-local hosts, and a missing tool allowlist, as text in the repo. It does not run these 2,500 trials, and it does not measure whether the declared setting is what the agent process enforced. The bench's lesson for that scanner is to score the declaration and to keep task utility in the same sentence as attack success.

HarnessSecurity-Bench — arXiv 2026, arxiv.org/abs/2610.07639

SHarD — controls you can hand to a team

SHarD asks whether off-the-shelf controls can be packaged into one harness and installed for a team, instead of configured agent by agent. Built on Pi, it carried three: an OS sandbox (nono, kernel enforcement for the session, children included), skill scanning before a skill is loaded, and tool restriction on the bash tool. A content-protection control was dropped. During skill scanning it treated the scanner itself as an injection and broke the workflow. The adjusted score excludes those dropped tests so the comparison stays on the controls that shipped.

On a 23-test suite drawn from the OWASP Top 10 for Agentic Applications, the secured agents score: Claude Code 87.0% raw and 100% adjusted, Codex 69.6% and 75%, SHarD 78.3% and 100%. The adjusted 100% is the controlled categories, not every test in the suite. SHarD matched Claude Code on that adjusted score and did not regress the baseline Pi harness on the categories they re-ran. Skill scanning took 1m49s to 8m13s, against 18–30 seconds uncontrolled. The demo deny-list is a harness hook, and the paper is explicit that a call which never goes through the bash tool is outside it. The sandbox is the boundary they treat as structurally enforced. A pass that depended on the model noticing the scanner did not reproduce. On the unsandboxed baseline, the agent reached another agent's configuration on the same machine. They name the OS sandbox as the control that confines that reach.

Paniolo does not distribute a sandbox, a skill scanner, or a bash deny-list. The 100% is SHarD's adjusted score on their suite.

SHarD — arXiv 2026, arxiv.org/abs/2607.25890

SafeHarness — four layers, one threat model

SafeHarness places defenses inside the agent lifecycle: inform, verify, constrain, correct. The threat model is an adversary who controls external channels — pages, tool results, memory — and does not control the weights or the harness code. Against an unprotected baseline, across their configurations, they report about 38% lower unsafe behavior and about 42% lower attack success, with core task utility preserved and the interception cost tunable. Provenance tags on context are the layer that maps onto a wiki: a fact that entered from a tool keeps its origin. Paniolo's pages require a sources: path under raw/. That is attribution for a human and a linter. It is not their verifier.

SafeHarness — arXiv 2026, arxiv.org/abs/2604.13630
Untrusted writes

A small write
that lasts

AgentPoison — memory that outlives the turn

AgentPoison shows that unverified entries in long-term memory or a retrieval store can steer later behavior, at a poison rate below 0.1%. The poisoned items look like ordinary memory. A knowledge base that spot-checks a sample will miss a handful of writes. The scan rule for this class is memory-write provenance: a harness that copies untrusted tool or web output into durable memory without a quarantine step. The rule checks the declaration. It does not inspect a live memory store, and the paper's rate is not a Paniolo measurement.

AgentPoison — arXiv 2024, arxiv.org/abs/2407.12784

MCPTox — the description is the payload

MCPTox poisons tool metadata on live MCP servers. The poisoned tool is never executed. The instruction sits in the description the model reads while choosing a tool. The benchmark covers 45 real-world servers, 353 tools, and 1,312 cases. High attack success shows up across the agents they tried, and more capable models can be more susceptible. A permission that only gates execution still leaves the description in the prompt. The MCP server Paniolo exposes is scoped to one configured root. It does not review the description of every third-party tool an agent might load.

MCPTox — arXiv 2025, arxiv.org/abs/2508.14925
Diagnosis

Where, how,
and what followed

AgentDoG — a binary label misses the step

AgentDoG argues that safe or unsafe is the wrong output for an agent trajectory. A content filter misses risks that depend on tools and environment, and it misses a step that looks permitted on its own and is still unreasonable in the trace. The diagnostic splits three ways: where the risk entered, how it changed the agent's behavior, and what harm would follow. ATBench, built on that split, covers about 2,157 tools and 4,486 turn interactions. The released models are 4B, 7B, and 8B.

On R-Judge, AgentDoG-Qwen3-4B's F1 is 92.7. In the same table GPT-5.2's F1 is 91.8 and Gemini-3-Flash's is 95.3. The 4B model's accuracy cell on that row is also 91.8; the F1 is the comparison the paper draws. On ASSE-Safety, AgentDoG-Llama3.1-8B reaches 83.4 F1, against Gemini-3-Pro at 78.6. Specialized guard models in the table often keep high precision and very low recall. They miss unsafe intermediate steps and mark the trajectory safe. That is the diagnosis claim.

Scan does not classify a live trajectory and does not ship a guard model. The six posture checks — escalation protocol, resource caps, tool allowlists, untrusted-input boundaries, memory-write provenance, high-impact-action confirmation — verify that a protocol is declared. AgentDoG is the judgment layer. The taxonomy is still why a finding should say where, how, and what, instead of a single unsafe bit. Credentials are a separate store: paniolo secrets keeps them in the OS vault and passes selected names to one child process.

AgentDoG — arXiv 2026, arxiv.org/abs/2601.18491