291 MB
410 files
Updated about 1 month ago
Name
Size
asian_doll
billy_dog
boy_funko_pop
bull
cat_statue
ceramic_head
chicken_bean_bag
colorful_teapot
dangling_child
docs
elephant_sphere
elephant_statue
espresso_cup
gengar_toy
gold_pineapple
green_doll
iverson_funko_pop
maeve_dog
minion_toy
my_cat
rabbit_toy
red_chicken
red_piggy_bank
robot_toy
running_shoes
sheep_pillow
sheep_plush
sheep_toy
skulls_mug
small_penguin
.gitattributes2.31 kB
xet
README.md1.98 kB
xet
README.md

MyVLM

Paper: https://arxiv.org/abs/2403.14599

Project Page: https://snap-research.github.io/MyVLM/

Code: https://github.com/snap-research/MyVLM

MyVLM Objects Dataset

Example images for each object in our constructed dataset.

As part of our MyVLM code release, we have also released our object dataset introduced in the paper. This contains 29 user-specific objects, each containing ~10 images and 5 corresponding personalized captions for each image.

Your data should be organized using the following structure:

data_root
├── <concept_name>
│   ├── <image1>.jpg
│   ├── <image2>.jpg
│   ├── ...
│   ├── captions.json (or captions_augmented.json)
│   └── additional_llava_vqa_data.json  (optional, used for personalized VQA using LLaVA, see next section).
└── <concept_name_2>

That is, the root directory should contain a sub-directory for each concept. Then, in each concept directory, you should have:

  1. the set of images we want to use either for training or inference.
  2. a json file containing the captions for each image, named captions.json or captions_augmented.json. This file should be in the following format:
{
    "<image1>.jpg": ["<caption1>", "<caption2>", ...],
    "<image2>.jpg": ["<caption1>", "<caption2>", ...],
    ...
}

That is, we have a dictionary mapping each image path to a list of target captions. As described in the paper, at each optimization step we will randomly sample a caption from this list to use as the target caption for the image.

License

This sample code is made available by Snap Inc. for non-commercial, academic purposes only.

Please see the full license here.

Total size
291 MB
Files
410
Last updated
Jul 15
Pre-warmed CDN
US EU US EU

Contributors