Spaces:
Running
Running
File size: 8,020 Bytes
2e175db | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 | # Dataset curation
This folder contains the scripts that produced the **commercially-licensed
training dataset** for Stage 2 fine-tuning (DONE, 2026-05).
Stage 3A extends this folder with one new generation script per
approved open-source generator (initially AuraFlow, SDXL, and SD 3.5
Medium; PixArt-Ξ£ is deferred pending license/intended-use clarification).
The rest of the pipeline
(`build_manifest.py`, `split.py`, `../precompute_embeddings.py`,
`../train_head.py`) is generator-agnostic and will be reused as-is β
the manifest format already supports multiple AI source labels.
See [`docs/plan.md`](../../docs/plan.md) Stage 3 for the full plan.
## What we're building
A dataset of roughly **100k images** total:
| Split | Source | Class | License |
|------------|---------------------------------------|-----------------|--------------------|
| ~50k real | Open Images V7 (sampled) | `authentic` | CC BY 2.0 |
| ~50k AI | Self-generated with Flux.1-schnell | `ai_generated` | Apache 2.0 weights β outputs are unrestricted |
Stage 2 trains the CLIP classifier head on this 2-class data. The `deepfake`
and `edited` classes will be added later via separate scripts.
## Workflow
1. **`fetch_open_images.py`** β downloads a stratified sample of Open Images
into `data/raw/real/`. Run once. Idempotent.
2. **`generate_flux_synthetic.py`** β generates AI images with Flux.1-schnell
into `data/raw/ai_generated/`. Run on a rented cloud GPU
(~$30 total at RunPod / Lambda Labs spot rates).
**`generate_sdxl_synthetic.py`** β generates AI images with SDXL into
`data/raw/ai_generated/sdxl/`. Use `--dry-run` locally before running on a
GPU box.
**`generate_sd35_synthetic.py`** β generates AI images with SD 3.5 Medium
into `data/raw/ai_generated/sd35-medium/`. This source is conditionally
approved under the Stability AI Community License; review `NOTICES.md`
before running it.
**`generate_auraflow_synthetic.py`** β generates AI images with AuraFlow
v0.3 into `data/raw/ai_generated/auraflow-v0.3/`. Use `--dry-run` locally
before running on a GPU box.
Stage 3A generator scripts should reuse `generation_utils.py` for prompt
loading, deterministic seeds, stable output keys, file hashes, image
dimensions, approved generator metadata, and manifest-row construction.
3. **`build_manifest.py`** β produces `data/manifest.csv` with one row per
image. Required columns are
`path,class,source,license,license_url,sha256`. Stage 3A adds optional
generator metadata columns:
`generator,model_family,model_id,prompt,seed,width,height,generation_params_json`.
Missing optional values are written as empty strings, and approved
AI-generator metadata is backfilled from known `source` values when older
fragments do not include the optional columns. You can pass fragments
explicitly with `--inputs`, or point `--input-dir` at `data/raw` to discover
all `*_manifest.csv` fragments recursively. This is the
**legal record** that protects you when you commercialise and the
source-of-truth for per-generator evaluation.
4. **`split.py`** β produces train/val/test splits stratified by class. Stage
3A adds optional generator-aware splitting with
`--stratify-by class-generator`, plus `--holdout-generator <name>` to write
a `heldout.csv` evaluation split while excluding that generator from
train/val/test.
After splitting, training itself lives outside this directory:
5. **`augment_images.py`** β creates deterministic Stage 3A robustness
augmentations from any manifest or split CSV. It writes new JPEG images plus
a manifest-compatible CSV with `original_path`, `augmentation`,
`augmentation_seed`, and `augmentation_params_json`. Run it on `train.csv`
for augmented training rows, or on `val.csv` / `test.csv` into separate
manifests such as `test_augmented.csv` for robustness evaluation while
keeping the clean splits unchanged.
6. **`../precompute_embeddings.py`** β encodes every image through frozen
CLIP once and caches the 512-d feature vectors. Runs on CPU (1-4 hours
on the dev laptop). One-shot per dataset version. By default it encodes
clean `train.csv`, `val.csv`, and `test.csv`. Stage 3A adds
`--train-augment-manifest train_augmented.csv` to append augmented rows to
`train.npz`, and `--extra-split test_augmented=test_augmented.csv` for
explicit augmented robustness-eval embeddings.
7. **`../train_head.py`** β trains the small classifier head on the cached
features. Seconds per epoch on CPU; iterate hyperparameters freely. Stage
3A reports overall metrics, per-source/per-generator/per-family metrics when
embedding metadata is available, and augmentation robustness metrics for
explicitly encoded augmented eval splits. Use `--report-out` to save the
JSON report.
8. **`../evaluate_head.py`** β evaluates a candidate checkpoint, optionally
against a Stage 2 baseline checkpoint, on any cached embedding split
(`test`, `heldout`, `test_augmented`, etc.). It writes a reproducible JSON
report with overall, per-generator, held-out-generator, uncertainty, and
augmentation robustness metrics.
9. **`run_stage3a_smoke.py`** β runs a tiny local fixture pipeline before
expensive GPU work. It builds fixture images/manifests, splits with an SDXL
holdout, creates train/test augmentations, writes deterministic mocked
512-d embeddings, trains a smoke head, and evaluates it against a constant
baseline:
```bash
python scripts/dataset/run_stage3a_smoke.py \
--work-dir /tmp/deepfakescanner-stage3a-smoke
```
This intentionally bypasses real CLIP precompute so the smoke test stays
fast and offline; use `../precompute_embeddings.py` for real dataset runs.
10. **`../run_stage3a_pipeline.py`** β orchestrates the full Stage 3A sequence
after the smoke test passes. It can plan the run, execute GPU generation,
build the manifest, split, augment, precompute embeddings, train the
candidate head, evaluate against the Stage 2 baseline, write a JSON
ship/no-ship recommendation, and optionally upload an accepted checkpoint
to the private HF Hub repo.
```bash
python scripts/run_stage3a_pipeline.py --dry-run --stop-after evaluate
python scripts/run_stage3a_pipeline.py --stop-after generate
python scripts/run_stage3a_pipeline.py \
--skip-generation \
--baseline data/checkpoints/head_v1.pt
```
Publishing is intentionally opt-in via `--publish-if-accepted` and requires
`HF_TOKEN`.
For a collaborator-facing GPU handoff, use
[`docs/stage3a-gpu-collaborator-guide.md`](../../docs/stage3a-gpu-collaborator-guide.md).
## Why each piece is licensed for commercial use
- **Open Images V7** β Google's dataset, all images are CC BY 2.0 (commercial
use allowed with attribution). Attribution lives in `manifest.csv`.
- **Flux.1-schnell** β released by Black Forest Labs under Apache 2.0. Outputs
are not restricted; you own them. (Note: `flux.1-dev` is non-commercial β
do NOT use it.)
- **Stage 3A approved additions** β SDXL, SD 3.5 Medium, and AuraFlow were
reviewed on 2026-05-15. See [`NOTICES.md`](../../NOTICES.md) for the exact
license status and constraints before generating any images.
- **NOT used here**: FaceForensics++, DFDC, Celeb-DF (research-only),
outputs from closed-API generators (Gemini, Midjourney, DALL-E, Grok,
Firefly β ToS restrictions or ambiguous terms), `flux.1-dev`
(non-commercial), and PixArt-Ξ£ until its model-card intended-use ambiguity
is clarified. See [`NOTICES.md`](../../NOTICES.md) for the full list and
[`docs/decisions.md`](../../docs/decisions.md) 2026-05-14 and 2026-05-15
entries for the reasoning.
## Running it
The scripts are designed to be run independently. See each script's docstring
for prerequisites and command-line flags.
|