File size: 6,007 Bytes
b4bedeb
 
0b2df65
 
 
b4bedeb
 
 
 
 
 
 
 
 
 
0b2df65
b4bedeb
 
 
0b2df65
b4bedeb
0b2df65
b4bedeb
0b2df65
b4bedeb
 
 
 
 
 
0b2df65
b4bedeb
554b005
 
b4bedeb
 
 
 
4379d2c
b4bedeb
0fe19ce
 
 
 
 
 
 
 
 
 
 
 
b4bedeb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0b2df65
b4bedeb
 
 
 
 
 
 
 
0b2df65
b4bedeb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
# DetectiveSAMv2

DetectiveSAMv2 is an inference-only image forgery localization bundle built around SAM2. This repo is v2-only: the code path, config sidecar, tests, and demo commands all target the `detective_sam_v2` checkpoint.

The full runnable bundle with model weights is hosted on Hugging Face:

- https://huggingface.co/Gertlek/DetectiveSAMv2

## What is bundled

- Inference checkpoint configs under `checkpoints/`
- SAM2 config under `sam2configs/`
- Poster demo pairs under `demo/cocoglide/`, `demo/flux_test/`, and `demo/qwen_test/`
- A drop-in single-image slot at `demo/user_image/demo_input.png`

Built-in checkpoint alias:

- `detective_sam_v2`

The `detective_sam_v2` alias expects `checkpoints/detective_sam_v2.pth` and the bundled YAML sidecar.

## DetectiveSAMv2 Release

This release packages the v2 architecture and benchmarked weights without the old DetectiveSAMv1 compatibility branches.

Main changes:

- Evidence features built from unadapted target features, raw perturbation-stream features, and adapted target deltas.
- Spatial cross-attention feature adapters for flexible information sharing across target and perturbation streams.
- Transformer evidence mask adapter for direct prompt-mask prediction.
- V2-only adapter implementation in `detectivesam_inference/models/adapters.py`.
- Three perturbation streams by default: Gaussian blur, JPEG compression, and Gaussian noise.
- Training scaled beyond the original SIDA/MagicBrush mix with PicoBanana 10k clean, UltraEdit 10k, and AutoEditForge-generated QWEN/FLUX train data designed to promote generalization to new editors.
- Validation expanded across trained-on and held-out modern edit benchmarks, including QWEN-Bench, FLUX-Bench, CoCoGLIDE, AutoSplice, and NanoBanana.
- JSON and YAML checkpoint sidecar support, including training checkpoint containers with `model` or `ema` keys.

### Benchmark Scores

Scores below are IoU / F1 in percent for the 10-epoch v2 checkpoint `dsamv2_best6info_10ep_20260623_060047/best_model.pth`. Previous Detective SAM values are measured on the same validation setup. The MagicBrush row uses the MagicBrush test split, not the higher `full_magicbrush_val` artifact.

| Dataset | DetectiveSAMv2 | Previous Detective SAM | Delta |
| --- | ---: | ---: | ---: |
| MagicBrush test | 48.83 / 61.01 | 43.98 / 56.46 | +4.85 / +4.55 |
| SIDA test | 54.95 / 65.24 | 50.22 / 60.49 | +4.73 / +4.76 |
| FLUX-Bench | 44.26 / 56.17 | 18.61 / 25.77 | +25.65 / +30.39 |
| QWEN-Bench | 45.40 / 56.95 | 19.62 / 27.15 | +25.78 / +29.80 |
| CoCoGLIDE | 46.63 / 58.33 | 42.79 / 53.43 | +3.84 / +4.90 |
| AutoSplice | 57.49 / 70.53 | 47.85 / 60.32 | +9.64 / +10.21 |
| NanoBanana | 33.04 / 45.97 | 25.89 / 35.74 | +7.15 / +10.23 |
| Overall | 49.55 / 61.39 | 38.77 / 49.14 | +10.78 / +12.26 |

QWEN-Bench and FLUX-Bench are included because the training mix includes corresponding AutoEditForge-generated QWEN and FLUX datasets.

## Setup

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

Then download the large weights from Hugging Face:

```bash
pip install -U huggingface_hub
hf download Gertlek/DetectiveSAMv2 \
  checkpoints/detective_sam_v2.pth \
  sam2configs/sam2.1_hiera_base_plus.pt \
  --local-dir .
```

The expected checkpoint paths are also documented in `checkpoints/README.md` and `sam2configs/README.md`.

## Hugging Face Usage

For the simplest setup, clone the Hugging Face repo directly:

```bash
git lfs install
git clone https://huggingface.co/Gertlek/DetectiveSAMv2
cd DetectiveSAMv2
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m detectivesam_inference.predict \
  --output-dir outputs/poster_baseline
```

## Poster Demo Flows

### 1. Live single-image demo

Place your image at `demo/user_image/demo_input.png`, then run:

```bash
python -m detectivesam_inference.predict \
  --output-dir outputs/poster_user_image
```

In this mode the CLI reuses the target image as its own source reference so the demo stays runnable with a single image.

### 2. Bundled CocoGLIDE example

If `demo/user_image/demo_input.png` is absent, the default `predict` command falls back to the bundled CocoGlide sample `banana_28809`.

```bash
python -m detectivesam_inference.predict \
  --output-dir outputs/poster_baseline
```

### 3. Bundled modern-edit examples

Flux example:

```bash
python -m detectivesam_inference.predict \
  --checkpoint detective_sam_v2 \
  --source demo/flux_test/source/548.png \
  --target demo/flux_test/target/548.png \
  --mask demo/flux_test/mask/548.png \
  --output-dir outputs/poster_flux
```

Qwen example:

```bash
python -m detectivesam_inference.predict \
  --checkpoint detective_sam_v2 \
  --source demo/qwen_test/source/166.png \
  --target demo/qwen_test/target/166.png \
  --mask demo/qwen_test/mask/166.png \
  --output-dir outputs/poster_qwen
```

### 4. Bundled CocoGlide subset sweep

Use this to evaluate the bundled banana and train CocoGlide demo pairs.

```bash
python -m detectivesam_inference.evaluate \
  --checkpoint detective_sam_v2 \
  --dataset-root demo/cocoglide \
  --output-dir outputs/poster_eval_cocoglide \
  --num-visualizations 2
```

## Outputs

Each `predict` run writes a compact set of visual artifacts plus a JSON summary:

- `<name>_comparison.png`
- `<name>_probability.png`
- `<name>_pred_mask.png`
- `<name>_pred_overlay.png`
- `<name>_summary.json`

If a ground-truth mask is provided, the run also saves:

- `<name>_gt_mask.png`
- `<name>_gt_overlay.png`

The `evaluate` command writes `summary.json` plus a few visualization examples under `visualizations/`.

## Notes

- The runtime selects `cuda` automatically when available and otherwise runs on CPU.
- Checkpoint settings come from YAML or JSON sidecars in `checkpoints/`; you only need the alias or checkpoint path.
- The Hugging Face repo bundles the v2 `.pth` checkpoint and SAM2 `.pt` weight file.
- This repo does not include training code or training-only dependencies.