scanner / NOTICES.md
github-actions[bot]
Deploy from GitHub 39b3777315c11d9c8bcd39ad7bf034f2a88a7379 (filtered: code + Dockerfile + README + NOTICES only)
2e175db
|
Raw
History Blame Contribute Delete
7.55 kB

Third-Party Notices

This project is built on commercially-licensed open-source components only. Every dependency below has been verified against its license terms.

Runtime dependencies

Component Version (β‰₯) License Use
Python 3.11 PSF Runtime
PyTorch (CPU build) 2.3 BSD-3-Clause Tensor / NN runtime
torchvision 0.18 BSD-3-Clause Image transforms
HuggingFace transformers 4.41 Apache-2.0 CLIP loading
HuggingFace accelerate 0.30 Apache-2.0 Inference helpers
OpenAI CLIP ViT-B/32 n/a MIT Frozen image-encoder weights
Pillow (PIL fork) 10.3 HPND Image decoding
NumPy 1.26 BSD-3-Clause Tensor ops
FastAPI 0.111 MIT Web framework
Starlette (transitive) BSD-3-Clause ASGI core
Uvicorn 0.29 BSD-3-Clause ASGI server
python-multipart 0.0.9 Apache-2.0 File-upload parsing
Pydantic 2.7 MIT Schema validation
c2pa-python 0.5 Apache-2.0 C2PA / Content-Credentials verification

Development / test dependencies

Component License Use
pytest MIT Test runner
httpx BSD-3-Clause Test client

Stage 2 dataset-curation dependencies (NOT shipped with the inference image)

These are only installed on the GPU machine that builds the dataset.

Component License Use
FiftyOne Apache-2.0 Open Images V7 download / sampling
diffusers (HuggingFace) Apache-2.0 Flux.1-schnell pipeline
black-forest-labs/FLUX.1-schnell Apache-2.0 Synthetic-image generator weights

Stage 3A candidate generator license review

Review date: 2026-05-15. No Stage 3A images have been generated yet; this table records which generators are approved for the initial multi-generator dataset.

Component License / terms Status Use
black-forest-labs/FLUX.1-schnell Apache-2.0 (HF model card) Approved Already-used rectified-flow baseline.
stabilityai/stable-diffusion-xl-base-1.0 CreativeML Open RAIL++-M (HF license) Approved with use-policy restrictions Stage 3A diffusion U-Net training source. Outputs may be used for detector training, subject to OpenRAIL use restrictions; do not use to train a competing generative model.
stabilityai/stable-diffusion-3.5-medium Stability AI Community License (HF model card, Stability license) Conditionally approved Stage 3A MMDiT / diffusion-transformer training source while Veridicate remains under the Community License revenue threshold and registers as required. Enterprise license required before use if annual revenue exceeds USD $1M. Outputs must not be used to create or improve a foundational generative AI model.
fal/AuraFlow-v0.3 Apache-2.0 (HF model card) Approved Stage 3A independent rectified-flow training source.
PixArt-alpha/PixArt-Sigma-XL-2-1024-MS HF model card labels weights as CreativeML Open RAIL++-M; project code is Apache-2.0, but the model card's direct-use section says research purposes only (HF model card, GitHub repo) Deferred / blocked pending clarification Do not use for Stage 3A training until the model-weight license and model-card intended-use language are clarified.

Datasets used at training time

Dataset License Notes
Open Images V7 CC BY 2.0 Real-image class. 50,000 images sampled. Attribution preserved in data/manifest.csv.
Self-generated (Flux schnell) Apache-2.0 (model) AI-image class. 50,000 images generated. We own the outputs; recorded in manifest with prompt + seed.

Distributed artifacts (downstream of training)

Artifact License Notes
Veridicate/scanner-head-v1 (private HF Hub model repo) Apache-2.0 Trained CLIP classifier head weights (~530 KB). Downloaded at container startup by ClipClassifier using the HF_TOKEN Space secret. Trained on the Open Images + Flux-schnell dataset above, so the output weights are unencumbered. Repo is private (commercial-IP reasons) but the license on the weights themselves is Apache-2.0.

Explicitly NOT used (research-only / non-commercial / unclear license)

These are deliberately excluded to preserve commercial usability:

  • FaceForensics++ β€” research/educational use only.
  • DFDC (Deepfake Detection Challenge) β€” research only.
  • Celeb-DF β€” research only.
  • flux.1-dev β€” non-commercial license. Use flux.1-schnell instead.
  • Outputs from closed-API image generators β€” including Google Gemini / Imagen 3 / "Nano Banana", OpenAI DALL-E, Midjourney, xAI Grok Imagine / Aurora, Adobe Firefly. Their terms of service either restrict using outputs as training data for downstream models or leave it ambiguous enough that, for a commercial detector, the safe reading is "not allowed without an explicit license." Stage 3 of the roadmap (see docs/plan.md) explicitly catches these generators by training on the open-source equivalents of their underlying architectures (diffusion U-Net, DiT, rectified flow, GAN) and learning the family fingerprint β€” without ever using their outputs as training data. Stage 5 may revisit the closed-API path with negotiated enterprise license agreements, but only once revenue justifies the legal and per-image API spend.
  • Any HuggingFace model whose license tag is cc-by-nc-*, non-commercial, or unspecified.

Updating this document

Add a row whenever a new dependency is introduced. If a license is unclear, flag it and resolve before merging β€” never assume.

For forward-looking license concerns (revenue thresholds, deferred items, in-flight investigations), see docs/licensing.md. That file is the live tracker; this file is the approved-and-shipped record.