umm / data /covt /README.md
Orangerl's picture
Add files using upload-large-folder tool
20b4034 verified
|
Raw
History Blame Contribute Delete
1.88 kB

Stage-2 CoVT data

This is the exact merged local dataset used by the current Stage-2 recipe:

  • 889,489 rows;
  • 756,894 unique files in the flat image directory;
  • every row has exactly one image plus a multi-turn conversation;
  • 3,387,330 conversation messages in total;
  • no T2I target images and no I2I source/target pairs.

Stage 2 therefore trains visual understanding/VQA language generation. The five frozen visual teachers create online supervision from each row's input image; they are not extra target columns stored in JSON.

Source labels visible in the merged IDs

The original merged JSON contains only id, image, and conversations; it does not contain a formal provenance column. The following counts are audited from the source labels encoded in id and should not be treated as a new upstream-license declaration.

Encoded source label Rows
TallyQA (tallyqa plus cauldron/tallyqa) 363,591
numeric/unlabeled IDs 99,797
image textualization 99,573
ALLaVA (laion plus vflan) 69,980
Cauldron CLEVR 69,995
ShareGPT4o 57,284
identity 29,193
Cauldron IconQA 27,302
Cauldron ST-VQA 17,242
Cauldron A-OKVQA 16,534
Cauldron Visual7W 14,361
SA segmentation subset 8,990
Cauldron Hateful Memes 8,495
ADE depth subset 5,000
Cauldron VSR 2,152

The counts sum to 889,489. Because 99,797 rows have numeric IDs only, exact upstream attribution for those rows cannot be reconstructed from this merged artifact alone.

Portable representation

data.json rewrites every absolute source-machine image path to a basename. The loader resolves those basenames against data/covt/images. The 64 tar files under shards/ materialize that directory; each tar has a neighboring SHA256 sidecar and is verified before extraction by scripts/bootstrap.sh.