Papers
arxiv:2609.35932

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Published on Sep 28
· Submitted by
Yan Zhan
on Sep 30
Authors:
,
,
,
,

Abstract

Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.

Community

Paper author Paper submitter

Prompt-injection defenses often treat a chat-template marker as ordinary text. This paper shows that the same visible bytes can carry very different authority depending on whether the tokenizer emits a reserved control token or ordinary subwords. Holding the decoded text fixed and controlling for token count, we isolate the learned reserved-token representation as the mechanism behind a large part of the attack gap: replacing reserved markers with subwords reduces InjecAgent attack success by 39–66 percentage points across three open-weight model families, and the effect transfers to multi-turn AgentDojo tasks. An embedding-level analysis shows that the nearest ordinary-token vector can restore the attack on Llama-3.1, while adaptive marker search finds such embedding neighbors in three of four families. We also audit tokenizer-side mitigations and show that many configurations leave tool-protocol markers exposed—the same channel agents use to read untrusted tool output.

📄 Code: https://github.com/Byte-Authority/Byte-Authority
🤗 Evaluation data: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation

Wow, this is a pretty cool and smart way to prevent a lot of prompt injection attacks as far as I can see. I like the creative thinking approaching the problem from the tokenizer side. That's neat af.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35932
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35932 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35932 in a Space README.md to link it from this page.

Collections including this paper 1