File size: 15,878 Bytes
c01b9cd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4e2e30e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
# PoC Tester Agent β€” Progress Log

> **Repository:** `uandersonricardo/projeto-talp1`
> **Branch:** `teste_agent`
> **Model used:** `google/gemini-3.1-flash-lite` (via OpenRouter)
> **Framework:** LangGraph + Foundry
> **Dataset:** [ASSERT-KTH/Proof-of-Patch](https://github.com/ASSERT-KTH/Proof-of-Patch) β€” 22 real-world smart contract vulnerabilities with verified patches

---

## Metric Definitions

| Metric | Formula | Meaning |
|---|---|---|
| **Reproducibility Rate** | `PoCs passing on vulnerable version / 22` | Agent generates a working exploit |
| **Specificity Rate** | `PoCs failing on patched version / reproducible PoCs` | Exploit is logically correct (patch stops it) |
| **Overall Ground Truth** | `reproducible AND specific / 22` | Both conditions satisfied |

> **Reference:** The PoCo paper (Andersson et al., arXiv:2511.02780) evaluated on the same dataset using GPT-4o.
> Cases #003 and #015 have "inconclusive patches" per Β§5.3.1 β€” the patch fixes the bug but the PoC still passes because it tests side-effects unaffected by the fix. This is a known dataset limitation.

---

## Benchmark History

### Run 0 β€” Baseline (before any improvements)
**Date:** 2026-06-07 (pre-session)
**Commit:** `d849e68`

| Metric | Value |
|---|---|
| Reproducibility | **27.3%** (6/22) |
| Specificity | **0.0%** (0/6) |
| Overall Ground Truth | **0.0%** |
| Avg Iterations | 8.18 |

**Reproducible:** 001, 003, 008, 051, 054, 091
**Main failure:** ~63% were COMPILER_ERROR β€” agent guessing wrong import paths with no context.

---

### Run 1 β€” Quick validation (5 cases, first improvements)
**Date:** 2026-06-08

| Metric | Value |
|---|---|
| Reproducibility | **20.0%** (1/5) |
| Specificity | **0.0%** |
| Avg Iterations | 8.40 |

---

### Run 2 β€” Quick validation (5 cases, BFS + minimal interface)
**Date:** 2026-06-08
**Key changes:** BFS test file selection, compileFailures counter, MINIMAL_INTERFACE escape hatch

| Metric | Value |
|---|---|
| Reproducibility | **80.0%** (4/5) |
| Specificity | **0.0%** |
| Avg Iterations | 7.00 |

**Notable:** Case 015: 10 iters β†’ failed became 6 iters β†’ success thanks to MINIMAL_INTERFACE.

---

### Run 3 β€” Full benchmark (22 cases)
**Date:** 2026-06-08
**Commit:** `1929f1b`

| Metric | Value |
|---|---|
| Reproducibility | **54.5%** (12/22) |
| Specificity | **0.0%** (0/12) |
| Overall Ground Truth | **0.0%** |
| Avg Iterations | 7.55 |

**Reproducible (12):** 001, 003, 008, 015, 032, 051, 054, 058, 066, 077, 091, 098
**Failed (10):** 009, 018, 020, 033, 039, 041, 042, 048, 049, 070
**Error (1):** 046 (forge not in PATH during setup)

**Root cause of specificity=0%:** Patch application was broken β€” `cp -rv patches/ID/*` copied
a repo-name subdirectory INTO tempPatchDir instead of overwriting the actual source files.
The "patched" version was still running the vulnerable code.

---

## Per-Case Results (Run 3)

| ID | Vuln Type | Project | Repro | Iter | Failure Root Cause |
|---|---|---|---|---|---|
| 001 | multicall | 2024-06-size | βœ… | 9 | Patch not applied (dir bug) |
| 003 | access control | 2023-07-pooltogether | βœ… | 5 | Inconclusive patch (paper Β§5.3.1) |
| 008 | logic error | 2023-09-centrifuge | βœ… | 2 | Patch not applied (dir bug) |
| 009 | logic error | 2023-10-caviar | ❌ | 10 | `lib/caviar/lib/oracle` missing |
| 015 | access control | 2023-07-pooltogether | βœ… | 5 | Inconclusive patch (paper Β§5.3.1) |
| 018 | flash loan | 2023-10-caviar | ❌ | 10 | `lib/caviar/lib/oracle` missing |
| 020 | denial of service | 2023-10-dopex | ❌ | 10 | `node_modules/@openzeppelin` missing |
| 032 | access control | 2022-06-putty | βœ… | 4 | Patch not applied (dir bug) |
| 033 | logic error | 2023-10-caviar | ❌ | 10 | `lib/caviar/lib/oracle` missing |
| 039 | unchecked calls | 2024-03-axis-finance | ❌ | 10 | Compiled OK but logic reverted |
| 041 | reentrancy | 2024-03-axis-finance | ❌ | 10 | Compiled OK but logic reverted |
| 042 | access control | 2023-10-cap | ❌ | 10 | `node_modules/@openzeppelin-upgradeable` missing |
| 046 | n/a | n/a | ❌ | β€” | forge not in PATH (setup script error) |
| 048 | reentrancy | 2023-10-caviar | ❌ | 10 | `lib/caviar/lib/oracle` missing |
| 049 | access control | 2024-01-salty | ❌ | 10 | `test/lib/UserFactory.sol` missing |
| 051 | logic error | 2023-11-panoptic | βœ… | 2 | Patch not applied (dir bug) |
| 054 | logic error | 2024-02-wise-lending | βœ… | 7 | Patch not applied (dir bug) |
| 058 | logic error | 2024-04-renzo | βœ… | 5 | Patch not applied (dir bug) |
| 066 | unchecked calls | 2024-05-munchables | βœ… | 8 | Patch not applied (dir bug) |
| 070 | reentrancy | 2024-08-ph | ❌ | 10 | `node_modules/@prb/test` missing |
| 077 | reentrancy | 2024-07-templegold | βœ… | 8 | Patch not applied (dir bug) |
| 091 | logic error | 2024-08-basin | βœ… | 9 | Patch not applied (dir bug) |
| 098 | reentrancy | 2022-05-cally | βœ… | 2 | Patch not applied (dir bug) |

---

## Improvements Implemented

### 1. `projectContextExtractor.ts` (NEW)
**Problem:** LLM was guessing import paths β†’ 63% COMPILER_ERROR failures.
**Fix:** Before generating PoC, reads `remappings.txt`, `foundry.toml`, and the most import-rich
`.t.sol` test file (BFS across all subdirs, picks file with most import lines).

### 2. `compileFailures` Counter + MINIMAL_INTERFACE Escape Hatch
**Problem:** After 10 failed compile attempts, LLM stuck in loop on wrong imports.
**Fix:** Counter increments per compile failure. After 3 consecutive β†’ switches to
`POC_MINIMAL_INTERFACE_PROMPT` which forbids all external imports and uses inline interfaces.

### 3. Error Context in Fix Prompt
**Problem:** LLM only saw "Import path is WRONG" β€” not which file.
**Fix:** `analyzeFoundryLog` extracts actual error lines (file path + line number) and surfaces
them at the top of the fix prompt.

### 4. Patch Diff in Analysis Prompt
**Problem:** LLM generating generic assertions passing on both vulnerable and patched versions.
**Fix:** Patch diff (unified diff of vulnerable vs patched contract) included in
`analyzeVulnerabilityNode` with instruction: "assertion must PASS on vulnerable, FAIL on patched."

### 5. `dependencyStubber.ts` (NEW)
**Problem:** 7 cases fail because `lib/caviar/lib/oracle/...` or `node_modules/@openzeppelin/...`
are absent from the sandbox.
**Fix:** Runs `forge build` probe before generation, detects all "Source not found" errors,
creates minimal stub contracts at those exact paths.

### 6. `applyPatchSmart` β€” Smart Patch Application (CRITICAL for specificity)
**Problem:** `cp -rv patches/ID/*` was copying a repo-name subdirectory INTO `tempPatchDir`.
The "patched" test was running the vulnerable code. This is why specificity was always 0%.
**Fix:** For each `.sol` in the patch dir, strips 1, 2, then 3 path prefix levels to find
matching file in `tempPatchDir` and copies it correctly.

```
Patch file: patches/003/2023-07-pooltogether/vault/src/Vault.sol
tempPatchDir: copy of findings/003/2023-07-pooltogether/vault/

strip=1: vault/src/Vault.sol β†’ NOT in tempPatchDir
strip=2: src/Vault.sol β†’ EXISTS βœ… β†’ copy applied
```

### 7. `computePatchDiff` β€” Correct Diff Calculation
**Problem:** Previous diff command tried a hardcoded path that didn't match the nested structure.
**Fix:** Uses same strip-depth logic as `applyPatchSmart` to find the right file pair.

### 8. Improved Prompts
- `ANALYZE_VULNERABILITY_PROMPT`: asks for specific assertion that fails after patching
- `POC_COMPILE_FIX_PROMPT`: import resolution hierarchy (remappings β†’ existing test β†’ inline)
- `POC_TEST_FIX_PROMPT`: added reentrancy `receive()` callback pattern + unchecked return pattern
- `POC_MINIMAL_INTERFACE_PROMPT` (NEW): full template for zero-external-import strategy

---

## Agent Architecture

```
VulnerabilityReport
       β”‚
       β–Ό
  oracleNode
  β”œβ”€β”€ generateLocalScaffold()
  β”œβ”€β”€ analyzeSolidityFile()           ← contract API extraction
  β”œβ”€β”€ extractProjectContext()          ← remappings + best .t.sol (BFS)
  └── createMissingDependencyStubs()   ← stubs missing lib/node_modules
       β”‚
       β–Ό
  analyzeVulnerabilityNode            ← LLM: root cause + specific assertion
  [ANALYZE_VULNERABILITY_PROMPT + patch diff]
       β”‚
       β–Ό
  generatePoCNode ←──────────────────────────┐
  β”œβ”€β”€ INITIAL: [POC_INITIAL_PROMPT]           β”‚
  β”œβ”€β”€ FIX_COMPILE (failures < 3):             β”‚
  β”‚   [POC_COMPILE_FIX_PROMPT]               β”‚
  β”œβ”€β”€ MINIMAL_INTERFACE (failures >= 3):      β”‚
  β”‚   [POC_MINIMAL_INTERFACE_PROMPT]          β”‚
  └── FIX_LOGIC: [POC_TEST_FIX_PROMPT]       β”‚
       β”‚                                      β”‚
       β–Ό                                      β”‚
  runFoundryNode β†’ analyzeFoundryLog          β”‚
       β”‚                                      β”‚
       β”œβ”€β”€ success ──────────────────── END   β”‚
       β”‚                                      β”‚
       └── failed + iters < 10 β†’ reflectNode β”€β”˜
```

---

## Next Steps

### P0 β€” Run benchmark with patch fix + dep stubs (Run 4)
Expected: Reproducibility ~68%, Specificity ~30%, Ground Truth ~20%

### P1 β€” Add `refineSpecificityNode`
If PoC passes on both versions, run a refinement step:
- Show: current PoC + patch diff
- Ask: "Make the assertion target exactly what the patch changes"

### P2 β€” Docker update
The Dockerfile builds TypeScript, so new files are included automatically.
Fix needed: `setup-sandbox.sh` may not find `forge` in Docker because foundryup
sets PATH in `~/.bashrc` (not sourced in non-interactive shells). Fix: add
`source ~/.foundry/env` or `export PATH="$HOME/.foundry/bin:$PATH"` explicitly.

### P3 β€” Integration check
The tester agent's public API (`VulnerabilityReport β†’ { status, solidityCode }`) is unchanged.
The server.ts route calling `runPoCGenerator()` still works.
Docker image needs rebuild after code changes.

---

## Expected Run 4 Results

| Metric | Run 3 | Target Run 4 |
|---|---|---|
| Reproducibility | 54.5% | ~68% |
| Specificity | 0.0% | ~30% |
| Overall Ground Truth | 0.0% | ~20% |
| Avg Iterations | 7.55 | <7.0 |


---

## Production vs Benchmark Gap Analysis

### What the production pipeline actually sends to the tester

```
User request
    β”‚
    β–Ό Coder Agent
coderResult.contract  ← a single Solidity string (the generated contract)
    β”‚
    β–Ό Auditor Agent (receives repoPath with Contract.sol + README.md)
auditorResult.findings[0]  ← ONE finding
    β”‚
    β–Ό mapFindingToReport()  ← the ONLY bridge between auditor and tester
```

**`mapFindingToReport` currently passes to the tester:**
| Field | Source | Used by tester? |
|---|---|---|
| `id` | derived from title | βœ… identifier only |
| `severity` | `finding.severity` | label only |
| `type` | `finding.type` | βœ… in analysis prompt |
| `title` | `finding.title` | βœ… in analysis prompt |
| `description` | `finding.description` | βœ… in analysis prompt |
| `affectedContract.sourceCode` | `coderResult.contract` | βœ… shown to LLM |
| `affectedContract.name` | extracted from path | βœ… |
| `attackVector` | `exploitablePaths[0]` | βœ… in analysis prompt |
| `exploitablePaths` | `judgeReview.exploitablePaths` | βœ… passed |
| `codeSnippet` | `finding.codeSnippet` | βœ… passed |
| `location` | `finding.location` | βœ… passed |

**What the auditor produces but the tester NEVER receives:**
| Auditor field | Value | Why it would help the tester |
|---|---|---|
| `finding.recommendation` | Human-readable fix suggestion | Tells tester what the PATCH would look like β†’ key for specific assertions |
| `finding.judgeReview.review` | Judge's analysis of exploitability | More precise attack reasoning than just description |
| `finding.judgeReview.confidence` | 0-100 confidence score | Tester could skip low-confidence findings |
| `auditorResult.repoContext` | Full structured protocol context | Gives tester knowledge of cross-contract interactions |
| `auditorResult.fileTree` | Directory tree of the repo | Helps tester find the right imports |
| All other `.sol` files | Source of ALL contracts | Tester only gets ONE contract; misses dependencies |

**Missing in production but present in benchmark:**
| Field | Benchmark | Production |
|---|---|---|
| `customSandboxDir` | βœ… real project folder | ❌ NOT SET (uses generic /tmp/poc-sandbox) |
| `referenceTestCode` | βœ… real test files | ❌ NOT SET |
| `patchDiff` | βœ… computed from dataset | ❌ N/A (no patch in production) |

**The critical gap:** In production, `customSandboxDir` is null/undefined, so:
- BFS test file selection β†’ SKIPPED
- projectContextExtractor β†’ SKIPPED  
- dependencyStubber β†’ SKIPPED
- Test file cleanup β†’ SKIPPED
- The tester runs in the generic `/tmp/poc-sandbox` with NO project context

All the improvements that boosted benchmark from 27% β†’ 54% are benchmark-only.

---

## Model vs Agent Quality Plateau

**How to tell them apart:**

| Symptom | Model limit | Agent limit |
|---|---|---|
| Correct assertion but setup wrong | | βœ… Agent can fix |
| Wrong assertion (generic, too broad) | βœ… Model limit | Could improve with better prompting |
| Compile errors even with MINIMAL_INTERFACE | | βœ… Agent can fix |
| Passes vulnerable but also passes patched | βœ… Model semantics | Partially agent (patch context) |
| Correct overall but random failures across runs | βœ… Temperature/stochastic | |

**Current evidence points:**
- gemini-3.1-flash-lite is a very small/cheap model β†’ likely hitting model ceiling for complex semantic reasoning
- The compile errors are 100% agent-fixable (and we mostly did)  
- Specificity failures: partly agent bug (patch not applied) + partly model (generic assertions)

**Test: try 3 cases with a stronger model to see the ceiling:**
```bash
OPENROUTER_MODEL="google/gemini-2.0-flash" BENCHMARK_LIMIT=3 npx tsx src/benchmark/runTesterBenchmark.ts
```
If specificity jumps to >50% with the stronger model, it's the model. If not, it's the agent prompting.

---

## Next Actions (Priority Order)

### 1. Fix production gap β€” `mapFindingToReport` (HIGH VALUE, LOW EFFORT)
Add the missing rich context from the auditor to what the tester receives:

```typescript
export function mapFindingToReport(finding: any, sourceCode: string, 
                                    repoContext?: string): VulnerabilityReport {
  return {
    // ... existing fields
    description: [
      finding.description,
      finding.recommendation ? `\nFix recommendation: ${finding.recommendation}` : "",
      finding.judgeReview?.review ? `\nJudge analysis: ${finding.judgeReview.review}` : "",
      repoContext ? `\nProtocol context: ${repoContext.slice(0, 1000)}` : "",
    ].filter(Boolean).join("\n"),
    // The recommendation tells the tester what should NOT work after the fix
    // which is the key for writing a specific assertion
  };
}
```

### 2. Pass `repoPath` as `customSandboxDir` in production
In server.ts, pass `outputDir` (where Contract.sol was written) as `customSandboxDir`:
```typescript
const report = mapFindingToReport(auditorResult.findings[0], coderResult.contract, 
                                   auditorResult.repoContext);
report.customSandboxDir = outputDir;  // ← enables all context improvements
```
Since production uses a single Contract.sol with no remappings or test files, the context
extractor will find no remappings (graceful fallback), and the dep stubber will run a 
probe but find nothing missing (also fine).

### 3. Lower MAX_ITERATIONS from 10 to 6
Credits are limited. 10 iterations is too many for flash-lite which repeats itself after ~5.

### 4. Run full benchmark to measure Run 4
After the test file cleanup fix + patch fix are confirmed working.