Papers
arxiv:2609.38807

PatchHolmes: Agentic Patch Retrieval via Listwise Selection

Published on Sep 30
· Submitted by
David Yang
on Oct 1
Authors:
,
,
,

Abstract

Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.

Community

Paper author Paper submitter

Accepted at AACL-IJCNLP 2026, Main Conference.

Every software vulnerability needs to be paired with the commit that fixed it. Security advisories need that pairing. So do severity scores, affected-version trackers, and supply-chain scanners.

❌ In the GitHub Advisory Database and the National Vulnerability Database, 60% to 63% of entries do not have it.

The search is hard for three reasons:

❌ The haystack is large. A repository can carry up to 1.4 million commits, and 49% of these vulnerabilities are filed against repositories with more than 5,000.
❌ The candidates are long. The average commit diff in our corpus runs about 15,000 tokens. The text encoders that earlier methods use read the first 512.
❌ The words do not match. The vulnerability report names the symptom, "buffer overflow". The commit that fixes it names the bug, "out-of-bounds read".

🔍 The failure that started this paper.

The previous best agentic system scores each of 10 candidates in its own separate call: yes or no, plus a confidence.

On CVE-2015-9251, a cross-site scripting bug in jQuery, it marks all ten "yes" at confidence 5. Nothing distinguishes them, so the rank-1 pick is whatever the list happened to put first. That was an unrelated refactor of the same code path.

✅ Our fix: let the agent see the whole list.

PatchHolmes reads the top 100 candidates at once, then opens individual commits through four budgeted tools and submits a single answer. The search narrows in four steps: up to 1.4 million commits in the repository, 100 candidates after the first stage, 5 commits the agent actually opens, 1 answer submitted. On that jQuery case it reads five commits, in list order 1, 2, 9, 4, 3, and picks the one that blocks automatic execution of JavaScript responses.

📊 Results. Recall@1 is the share of vulnerabilities where the single submitted commit is the correct one, so higher is better.

✅ PatchHolmes 59.95%, against 34.61% for the pointwise system and 28.55% for a general retrieve-and-reason baseline
✅ On identical candidates, the agent lifts Recall@1 from 32.63% to 59.95%
✅ The same agent, moved unchanged to another benchmark's candidate pool, lifts Recall@1 from 24.28% to 39.86%
✅ Swapping the language model inside one family moves Recall@1 by under 1 point
✅ About 96,000 input tokens, 8 tool calls, and 5 commits read per vulnerability

The whole system runs on a frozen open-weight model over a local Git clone. No fine-tuning and no paid search API, because both cost more as the backlog grows.

🙏 Joint work with Yingming Zhou, Jiangrui Zheng, Shudong Hao, and Xueqing Liu at Stevens Institute of Technology.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.38807 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.38807 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.38807 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.