Buckets:
| # prelinger-films | |
| Mirror of the explicitly public-domain subset of the [Prelinger Archives](https://archive.org/details/prelinger) | |
| (Internet Archive), prepared for batch video captioning with | |
| [uv-scripts/video](https://huggingface.co/datasets/uv-scripts/video) (Marlin-2B on HF Jobs). | |
| **Selection**: `collection:prelinger AND mediatype:movies AND licenseurl:*publicdomain*` | |
| — 1,876 items, ~381 hours. Every item carries an explicit Creative Commons public | |
| domain mark in its IA metadata; see each film's sidecar for the per-item rights basis. | |
| Context: the archive states ~65% of its holdings are US public domain, and materials | |
| hosted on IA carry a standing public-domain dedication supporting commercial and | |
| non-commercial reuse. PD status is a US determination. | |
| ## Layout | |
| - `films/{identifier}.mp4` — best available MP4 derivative per item (512kb preferred) | |
| - `meta/{identifier}.json` — per-film sidecar: `identifier`, `licenseurl`, `title`, | |
| `date`, `runtime`, source URL, fetch timestamp. Sidecar-exists == film fully copied. | |
| - `chunks/{identifier}/c%04d.mp4` — ~60s stream-copy segments (keyframe-snapped) | |
| - `chunkmeta/{identifier}.json` — actual segment boundary times (offsets for | |
| timestamp remapping are these recorded values, not assumed 60s multiples) | |
| - `captions/` — saturate output: resumable parquet, one row per chunk | |
| (`scene`, `caption`, `events` JSON with global-time `<start – end>` events) | |
| Built 2026-07-30 with HF Jobs; pipeline scripts archived in | |
| `hf://buckets/davanstrien/prelinger-sample/scripts/`. | |
Xet Storage Details
- Size:
- 1.55 kB
- Xet hash:
- 47db8ff1cc74cb3dcb13f94840aaea7410d15749442cb319a68f74c7a715009b
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.