Papers
arxiv:2609.34518

SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

Published on Sep 28
· Submitted by
Xinjie Shen
on Sep 30
Authors:
,
,
,

Abstract

Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at https://github.com/EverywhereSafety/SEAD.

Community

Paper submitter

SEAD studies agent safety through the state changes caused by tool use. An action that appears harmless in isolation can become harmful after earlier steps change the environment, while the visible conversation may not reveal that state. We formulate attack and defense as partially observed state control: DART uses execution feedback to guide adaptive trajectory search, and SAGE investigates relevant environment state through read-only queries before allowing or blocking a proposed action. Together, they study how to intercept harmful transitions while preserving legitimate progress.
image

Every agent I've shipped logs its tool calls in beautiful detail and still can't tell me what the filesystem looks like right now, or which row got flipped. The transcript is narration, not state — and that's why prompt-level defenses are structurally blind here: the harmful step looks routine in isolation, so nothing in the text of the call reads as anomalous. What actually catches it is a per-step state journal diffed before each action, so the agent's own account stops being the source of truth. The case I'd want to see handled is the idempotent-looking write — a second write that changes nothing visible but flips a flag — because that's what kills naive diffing. And a full snapshot per step is cheap on a repo and brutal on a live database, so I'd expect the practical version to be a declared side-effect contract per tool rather than a general observer.

·

Cool point. My main concern with post-execution diffing is that the damage may already be done, for example rm -rf /. Sandboxing everything first could avoid that, but may be expensive.

That is why SAGE acts before execution. It selectively queries relevant state and then PASS/BLOCKs the pending action. We also released the SAGE-4B checkpoint and inference code.

I really like the side-effect-contract idea for making this practical.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.34518
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.34518 in a Space README.md to link it from this page.

Collections including this paper 1