Spaces:
Runtime error
Runtime error
File size: 15,878 Bytes
c01b9cd 4e2e30e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 | # PoC Tester Agent β Progress Log
> **Repository:** `uandersonricardo/projeto-talp1`
> **Branch:** `teste_agent`
> **Model used:** `google/gemini-3.1-flash-lite` (via OpenRouter)
> **Framework:** LangGraph + Foundry
> **Dataset:** [ASSERT-KTH/Proof-of-Patch](https://github.com/ASSERT-KTH/Proof-of-Patch) β 22 real-world smart contract vulnerabilities with verified patches
---
## Metric Definitions
| Metric | Formula | Meaning |
|---|---|---|
| **Reproducibility Rate** | `PoCs passing on vulnerable version / 22` | Agent generates a working exploit |
| **Specificity Rate** | `PoCs failing on patched version / reproducible PoCs` | Exploit is logically correct (patch stops it) |
| **Overall Ground Truth** | `reproducible AND specific / 22` | Both conditions satisfied |
> **Reference:** The PoCo paper (Andersson et al., arXiv:2511.02780) evaluated on the same dataset using GPT-4o.
> Cases #003 and #015 have "inconclusive patches" per Β§5.3.1 β the patch fixes the bug but the PoC still passes because it tests side-effects unaffected by the fix. This is a known dataset limitation.
---
## Benchmark History
### Run 0 β Baseline (before any improvements)
**Date:** 2026-06-07 (pre-session)
**Commit:** `d849e68`
| Metric | Value |
|---|---|
| Reproducibility | **27.3%** (6/22) |
| Specificity | **0.0%** (0/6) |
| Overall Ground Truth | **0.0%** |
| Avg Iterations | 8.18 |
**Reproducible:** 001, 003, 008, 051, 054, 091
**Main failure:** ~63% were COMPILER_ERROR β agent guessing wrong import paths with no context.
---
### Run 1 β Quick validation (5 cases, first improvements)
**Date:** 2026-06-08
| Metric | Value |
|---|---|
| Reproducibility | **20.0%** (1/5) |
| Specificity | **0.0%** |
| Avg Iterations | 8.40 |
---
### Run 2 β Quick validation (5 cases, BFS + minimal interface)
**Date:** 2026-06-08
**Key changes:** BFS test file selection, compileFailures counter, MINIMAL_INTERFACE escape hatch
| Metric | Value |
|---|---|
| Reproducibility | **80.0%** (4/5) |
| Specificity | **0.0%** |
| Avg Iterations | 7.00 |
**Notable:** Case 015: 10 iters β failed became 6 iters β success thanks to MINIMAL_INTERFACE.
---
### Run 3 β Full benchmark (22 cases)
**Date:** 2026-06-08
**Commit:** `1929f1b`
| Metric | Value |
|---|---|
| Reproducibility | **54.5%** (12/22) |
| Specificity | **0.0%** (0/12) |
| Overall Ground Truth | **0.0%** |
| Avg Iterations | 7.55 |
**Reproducible (12):** 001, 003, 008, 015, 032, 051, 054, 058, 066, 077, 091, 098
**Failed (10):** 009, 018, 020, 033, 039, 041, 042, 048, 049, 070
**Error (1):** 046 (forge not in PATH during setup)
**Root cause of specificity=0%:** Patch application was broken β `cp -rv patches/ID/*` copied
a repo-name subdirectory INTO tempPatchDir instead of overwriting the actual source files.
The "patched" version was still running the vulnerable code.
---
## Per-Case Results (Run 3)
| ID | Vuln Type | Project | Repro | Iter | Failure Root Cause |
|---|---|---|---|---|---|
| 001 | multicall | 2024-06-size | β
| 9 | Patch not applied (dir bug) |
| 003 | access control | 2023-07-pooltogether | β
| 5 | Inconclusive patch (paper Β§5.3.1) |
| 008 | logic error | 2023-09-centrifuge | β
| 2 | Patch not applied (dir bug) |
| 009 | logic error | 2023-10-caviar | β | 10 | `lib/caviar/lib/oracle` missing |
| 015 | access control | 2023-07-pooltogether | β
| 5 | Inconclusive patch (paper Β§5.3.1) |
| 018 | flash loan | 2023-10-caviar | β | 10 | `lib/caviar/lib/oracle` missing |
| 020 | denial of service | 2023-10-dopex | β | 10 | `node_modules/@openzeppelin` missing |
| 032 | access control | 2022-06-putty | β
| 4 | Patch not applied (dir bug) |
| 033 | logic error | 2023-10-caviar | β | 10 | `lib/caviar/lib/oracle` missing |
| 039 | unchecked calls | 2024-03-axis-finance | β | 10 | Compiled OK but logic reverted |
| 041 | reentrancy | 2024-03-axis-finance | β | 10 | Compiled OK but logic reverted |
| 042 | access control | 2023-10-cap | β | 10 | `node_modules/@openzeppelin-upgradeable` missing |
| 046 | n/a | n/a | β | β | forge not in PATH (setup script error) |
| 048 | reentrancy | 2023-10-caviar | β | 10 | `lib/caviar/lib/oracle` missing |
| 049 | access control | 2024-01-salty | β | 10 | `test/lib/UserFactory.sol` missing |
| 051 | logic error | 2023-11-panoptic | β
| 2 | Patch not applied (dir bug) |
| 054 | logic error | 2024-02-wise-lending | β
| 7 | Patch not applied (dir bug) |
| 058 | logic error | 2024-04-renzo | β
| 5 | Patch not applied (dir bug) |
| 066 | unchecked calls | 2024-05-munchables | β
| 8 | Patch not applied (dir bug) |
| 070 | reentrancy | 2024-08-ph | β | 10 | `node_modules/@prb/test` missing |
| 077 | reentrancy | 2024-07-templegold | β
| 8 | Patch not applied (dir bug) |
| 091 | logic error | 2024-08-basin | β
| 9 | Patch not applied (dir bug) |
| 098 | reentrancy | 2022-05-cally | β
| 2 | Patch not applied (dir bug) |
---
## Improvements Implemented
### 1. `projectContextExtractor.ts` (NEW)
**Problem:** LLM was guessing import paths β 63% COMPILER_ERROR failures.
**Fix:** Before generating PoC, reads `remappings.txt`, `foundry.toml`, and the most import-rich
`.t.sol` test file (BFS across all subdirs, picks file with most import lines).
### 2. `compileFailures` Counter + MINIMAL_INTERFACE Escape Hatch
**Problem:** After 10 failed compile attempts, LLM stuck in loop on wrong imports.
**Fix:** Counter increments per compile failure. After 3 consecutive β switches to
`POC_MINIMAL_INTERFACE_PROMPT` which forbids all external imports and uses inline interfaces.
### 3. Error Context in Fix Prompt
**Problem:** LLM only saw "Import path is WRONG" β not which file.
**Fix:** `analyzeFoundryLog` extracts actual error lines (file path + line number) and surfaces
them at the top of the fix prompt.
### 4. Patch Diff in Analysis Prompt
**Problem:** LLM generating generic assertions passing on both vulnerable and patched versions.
**Fix:** Patch diff (unified diff of vulnerable vs patched contract) included in
`analyzeVulnerabilityNode` with instruction: "assertion must PASS on vulnerable, FAIL on patched."
### 5. `dependencyStubber.ts` (NEW)
**Problem:** 7 cases fail because `lib/caviar/lib/oracle/...` or `node_modules/@openzeppelin/...`
are absent from the sandbox.
**Fix:** Runs `forge build` probe before generation, detects all "Source not found" errors,
creates minimal stub contracts at those exact paths.
### 6. `applyPatchSmart` β Smart Patch Application (CRITICAL for specificity)
**Problem:** `cp -rv patches/ID/*` was copying a repo-name subdirectory INTO `tempPatchDir`.
The "patched" test was running the vulnerable code. This is why specificity was always 0%.
**Fix:** For each `.sol` in the patch dir, strips 1, 2, then 3 path prefix levels to find
matching file in `tempPatchDir` and copies it correctly.
```
Patch file: patches/003/2023-07-pooltogether/vault/src/Vault.sol
tempPatchDir: copy of findings/003/2023-07-pooltogether/vault/
strip=1: vault/src/Vault.sol β NOT in tempPatchDir
strip=2: src/Vault.sol β EXISTS β
β copy applied
```
### 7. `computePatchDiff` β Correct Diff Calculation
**Problem:** Previous diff command tried a hardcoded path that didn't match the nested structure.
**Fix:** Uses same strip-depth logic as `applyPatchSmart` to find the right file pair.
### 8. Improved Prompts
- `ANALYZE_VULNERABILITY_PROMPT`: asks for specific assertion that fails after patching
- `POC_COMPILE_FIX_PROMPT`: import resolution hierarchy (remappings β existing test β inline)
- `POC_TEST_FIX_PROMPT`: added reentrancy `receive()` callback pattern + unchecked return pattern
- `POC_MINIMAL_INTERFACE_PROMPT` (NEW): full template for zero-external-import strategy
---
## Agent Architecture
```
VulnerabilityReport
β
βΌ
oracleNode
βββ generateLocalScaffold()
βββ analyzeSolidityFile() β contract API extraction
βββ extractProjectContext() β remappings + best .t.sol (BFS)
βββ createMissingDependencyStubs() β stubs missing lib/node_modules
β
βΌ
analyzeVulnerabilityNode β LLM: root cause + specific assertion
[ANALYZE_VULNERABILITY_PROMPT + patch diff]
β
βΌ
generatePoCNode ββββββββββββββββββββββββββββ
βββ INITIAL: [POC_INITIAL_PROMPT] β
βββ FIX_COMPILE (failures < 3): β
β [POC_COMPILE_FIX_PROMPT] β
βββ MINIMAL_INTERFACE (failures >= 3): β
β [POC_MINIMAL_INTERFACE_PROMPT] β
βββ FIX_LOGIC: [POC_TEST_FIX_PROMPT] β
β β
βΌ β
runFoundryNode β analyzeFoundryLog β
β β
βββ success ββββββββββββββββββββ END β
β β
βββ failed + iters < 10 β reflectNode ββ
```
---
## Next Steps
### P0 β Run benchmark with patch fix + dep stubs (Run 4)
Expected: Reproducibility ~68%, Specificity ~30%, Ground Truth ~20%
### P1 β Add `refineSpecificityNode`
If PoC passes on both versions, run a refinement step:
- Show: current PoC + patch diff
- Ask: "Make the assertion target exactly what the patch changes"
### P2 β Docker update
The Dockerfile builds TypeScript, so new files are included automatically.
Fix needed: `setup-sandbox.sh` may not find `forge` in Docker because foundryup
sets PATH in `~/.bashrc` (not sourced in non-interactive shells). Fix: add
`source ~/.foundry/env` or `export PATH="$HOME/.foundry/bin:$PATH"` explicitly.
### P3 β Integration check
The tester agent's public API (`VulnerabilityReport β { status, solidityCode }`) is unchanged.
The server.ts route calling `runPoCGenerator()` still works.
Docker image needs rebuild after code changes.
---
## Expected Run 4 Results
| Metric | Run 3 | Target Run 4 |
|---|---|---|
| Reproducibility | 54.5% | ~68% |
| Specificity | 0.0% | ~30% |
| Overall Ground Truth | 0.0% | ~20% |
| Avg Iterations | 7.55 | <7.0 |
---
## Production vs Benchmark Gap Analysis
### What the production pipeline actually sends to the tester
```
User request
β
βΌ Coder Agent
coderResult.contract β a single Solidity string (the generated contract)
β
βΌ Auditor Agent (receives repoPath with Contract.sol + README.md)
auditorResult.findings[0] β ONE finding
β
βΌ mapFindingToReport() β the ONLY bridge between auditor and tester
```
**`mapFindingToReport` currently passes to the tester:**
| Field | Source | Used by tester? |
|---|---|---|
| `id` | derived from title | β
identifier only |
| `severity` | `finding.severity` | label only |
| `type` | `finding.type` | β
in analysis prompt |
| `title` | `finding.title` | β
in analysis prompt |
| `description` | `finding.description` | β
in analysis prompt |
| `affectedContract.sourceCode` | `coderResult.contract` | β
shown to LLM |
| `affectedContract.name` | extracted from path | β
|
| `attackVector` | `exploitablePaths[0]` | β
in analysis prompt |
| `exploitablePaths` | `judgeReview.exploitablePaths` | β
passed |
| `codeSnippet` | `finding.codeSnippet` | β
passed |
| `location` | `finding.location` | β
passed |
**What the auditor produces but the tester NEVER receives:**
| Auditor field | Value | Why it would help the tester |
|---|---|---|
| `finding.recommendation` | Human-readable fix suggestion | Tells tester what the PATCH would look like β key for specific assertions |
| `finding.judgeReview.review` | Judge's analysis of exploitability | More precise attack reasoning than just description |
| `finding.judgeReview.confidence` | 0-100 confidence score | Tester could skip low-confidence findings |
| `auditorResult.repoContext` | Full structured protocol context | Gives tester knowledge of cross-contract interactions |
| `auditorResult.fileTree` | Directory tree of the repo | Helps tester find the right imports |
| All other `.sol` files | Source of ALL contracts | Tester only gets ONE contract; misses dependencies |
**Missing in production but present in benchmark:**
| Field | Benchmark | Production |
|---|---|---|
| `customSandboxDir` | β
real project folder | β NOT SET (uses generic /tmp/poc-sandbox) |
| `referenceTestCode` | β
real test files | β NOT SET |
| `patchDiff` | β
computed from dataset | β N/A (no patch in production) |
**The critical gap:** In production, `customSandboxDir` is null/undefined, so:
- BFS test file selection β SKIPPED
- projectContextExtractor β SKIPPED
- dependencyStubber β SKIPPED
- Test file cleanup β SKIPPED
- The tester runs in the generic `/tmp/poc-sandbox` with NO project context
All the improvements that boosted benchmark from 27% β 54% are benchmark-only.
---
## Model vs Agent Quality Plateau
**How to tell them apart:**
| Symptom | Model limit | Agent limit |
|---|---|---|
| Correct assertion but setup wrong | | β
Agent can fix |
| Wrong assertion (generic, too broad) | β
Model limit | Could improve with better prompting |
| Compile errors even with MINIMAL_INTERFACE | | β
Agent can fix |
| Passes vulnerable but also passes patched | β
Model semantics | Partially agent (patch context) |
| Correct overall but random failures across runs | β
Temperature/stochastic | |
**Current evidence points:**
- gemini-3.1-flash-lite is a very small/cheap model β likely hitting model ceiling for complex semantic reasoning
- The compile errors are 100% agent-fixable (and we mostly did)
- Specificity failures: partly agent bug (patch not applied) + partly model (generic assertions)
**Test: try 3 cases with a stronger model to see the ceiling:**
```bash
OPENROUTER_MODEL="google/gemini-2.0-flash" BENCHMARK_LIMIT=3 npx tsx src/benchmark/runTesterBenchmark.ts
```
If specificity jumps to >50% with the stronger model, it's the model. If not, it's the agent prompting.
---
## Next Actions (Priority Order)
### 1. Fix production gap β `mapFindingToReport` (HIGH VALUE, LOW EFFORT)
Add the missing rich context from the auditor to what the tester receives:
```typescript
export function mapFindingToReport(finding: any, sourceCode: string,
repoContext?: string): VulnerabilityReport {
return {
// ... existing fields
description: [
finding.description,
finding.recommendation ? `\nFix recommendation: ${finding.recommendation}` : "",
finding.judgeReview?.review ? `\nJudge analysis: ${finding.judgeReview.review}` : "",
repoContext ? `\nProtocol context: ${repoContext.slice(0, 1000)}` : "",
].filter(Boolean).join("\n"),
// The recommendation tells the tester what should NOT work after the fix
// which is the key for writing a specific assertion
};
}
```
### 2. Pass `repoPath` as `customSandboxDir` in production
In server.ts, pass `outputDir` (where Contract.sol was written) as `customSandboxDir`:
```typescript
const report = mapFindingToReport(auditorResult.findings[0], coderResult.contract,
auditorResult.repoContext);
report.customSandboxDir = outputDir; // β enables all context improvements
```
Since production uses a single Contract.sol with no remappings or test files, the context
extractor will find no remappings (graceful fallback), and the dep stubber will run a
probe but find nothing missing (also fine).
### 3. Lower MAX_ITERATIONS from 10 to 6
Credits are limited. 10 iterations is too many for flash-lite which repeats itself after ~5.
### 4. Run full benchmark to measure Run 4
After the test file cleanup fix + patch fix are confirmed working.
|