Add TOSC archive: 6 LoRA weights + 5 training sets + 9251 masked images + env specs
2c860a1 verified | # TOSC — 权重与训练数据归档 | |
| 基于 **masked image** 的 VLM 后训练工作:LoRA 权重 + 训练数据 + 复现所需的一切。 | |
| - 代码:<https://github.com/julyanghar/TOSC> | |
| - 本归档:<https://huggingface.co/julyanghar/TOSC> | |
| - TOSC benchmark 图片:<https://huggingface.co/julyanghar/tosc-model-weights> | |
| - 上游工作:[pspdada/SENTINEL](https://github.com/pspdada/SENTINEL) | |
| --- | |
| ## 目录结构 | |
| ``` | |
| . | |
| ├── README.md | |
| ├── manifest.json 机器可读账本:每个权重的数据/图片/超参 | |
| ├── weights/ 6 个 LoRA,每个附 train_config.md | |
| │ ├── llava-v1.5-13b-new-VG-masked-sentences/ | |
| │ ├── llava-v1.5-7b-new-VG-masked-sentences-split1-2/ | |
| │ ├── llava-v1.5-7b-VG-coco-sota-3-full/ | |
| │ ├── qwen2-vl-2b-masked-split1-2/ | |
| │ ├── qwen2-vl-7b-masked-sentences/ | |
| │ └── qwen2.5-vl-7b-masked-sentences/ | |
| ├── data/ 训练数据(图片路径已规范化) | |
| │ ├── <slug>.json 与 weights/<slug>/ 一一对应 | |
| │ └── _original_sources/ 未经改写的原始文件,用于核对 | |
| ├── images/ ⚠️ 按来源分两个子目录,不可拍平 | |
| │ ├── iterative_masking/ 5190 张 | |
| │ └── vg_coco_dino_only/ 4061 张 | |
| ├── wheels/ flash_attn 预编译轮子(191 MB) | |
| └── docs/ | |
| ├── ENVIRONMENTS.md 三个 conda 环境的复现指引 | |
| └── env-*.yml / *-pip.txt 三个环境的依赖清单 | |
| ``` | |
| --- | |
| ## ⚠️ 关于 `images/` 的两个子目录:**不要拍平** | |
| 两个来源目录里存在 **141 张同名但内容不同**的图片(实测 md5 全部不同)。数据 json 里的路径形如 `images/<source>/<file>`,那一层 `<source>` 是必需的——拍平成单目录会让 `llava-v1.5-7b-VG-coco-sota-3-full` 取到错误的图。 | |
| | 子目录 | 张数 | 原始来源 | 谁在用 | | |
| |---|---:|---|---| | |
| | `iterative_masking/` | 5190 | `SENTINEL_mask/vg_all_masked_images_from_iterative_masking` | 除 sota-3-full 外的 4 个数据集 | | |
| | `vg_coco_dino_only/` | 4061 | `results/VG-coco-collection/VG_coco_mask-dino-only/VG_coco_masked_image-LLAVA1.5-Hallucination/masked_image` | 仅 `llava-v1.5-7b-VG-coco-sota-3-full` | | |
| 所有图片已逐张 md5 校验,与权威来源一致(0 不一致)。 | |
| --- | |
| ## 6 个权重 | |
| | 权重 | Base | 训练数据 | 条数 | 图片 | 图片来源 | 框架 | | |
| |---|---|---|---:|---:|---|---| | |
| | `llava-v1.5-13b-new-VG-masked-sentences` | llava-v1.5-13b | ✅ | 10496 | 2964 | iterative_masking | SENTINEL 原生 | | |
| | `llava-v1.5-7b-new-VG-masked-sentences-split1-2` | llava-v1.5-7b | ✅ | 14207 | 3569 | iterative_masking | SENTINEL 原生 | | |
| | `llava-v1.5-7b-VG-coco-sota-3-full` | llava-v1.5-7b | ✅ 双来源互证 | 16384 | 4061 | **vg_coco_dino_only** | SENTINEL 原生 | | |
| | `qwen2-vl-2b-masked-split1-2` | Qwen2-VL-2B-Instruct | ✅ | 8918 | 1950 | iterative_masking | LLaMA-Factory | | |
| | `qwen2.5-vl-7b-masked-sentences` | Qwen2.5-VL-7B-Instruct | ✅ | 7686 | 2485 | iterative_masking | LLaMA-Factory | | |
| | `qwen2-vl-7b-masked-sentences` | Qwen2-VL-7B-Instruct | ⚠️ 未留存 | — | — | — | LLaMA-Factory | | |
| 全部 LoRA(r=128, alpha=256)。逐个权重的超参见 `weights/<slug>/train_config.md`。 | |
| **关于最后一个**:`qwen2-vl-7b-masked-sentences` 的训练数据没有留存。权重、超参(lr 2e-6、228 步、bs4 × ga8 × 2 卡)、训练日志都在,但训练集本身找不到了——LLaMA-Factory 不像 SENTINEL 原生框架那样会落 `dataset_logs`,且全盘扫过 1325 个 >2MB 的 json/jsonl 无匹配。要复现需按 pipeline 重新构造约 14592 条。 | |
| --- | |
| ## 数据格式 | |
| 两种,取决于训练框架。 | |
| **SENTINEL 原生(LLaVA 系)**: | |
| ```json | |
| { | |
| "id": "images/iterative_masking/2401137_masked_bicycle.jpg", | |
| "image": "images/iterative_masking/2401137_masked_bicycle.jpg", | |
| "image_path": "images/iterative_masking/2401137_masked_bicycle.jpg", | |
| "question": "What is this photo about? Please answer in great detail.", | |
| "context": "The image features a group of people ...", | |
| "y_win": "two women are looking back over two women's shoulders ...", | |
| "y_lose": "two women are riding bicycles ...", | |
| "type": "yw+_yl", | |
| "masked_image": "2401137_masked_bicycle.jpg" | |
| } | |
| ``` | |
| **LLaMA-Factory(Qwen 系)**: | |
| ```json | |
| { | |
| "instruction": "...", "context": "...", | |
| "chosen": "...", "rejected": "...", | |
| "images": ["images/iterative_masking/107930_masked_cup.jpg"] | |
| } | |
| ``` | |
| --- | |
| ## 快速开始 | |
| ```bash | |
| # 1) 建环境(见 docs/ENVIRONMENTS.md) | |
| conda env create -f docs/env-LLaMA-Factory-SENTINEL.yml | |
| # flash-attn 装不上时用现成的轮子: | |
| pip install wheels/flash_attn-2.7.3+cu12torch2.5cxx11abiFALSE-cp310-cp310-linux_x86_64.whl | |
| # 2) 把相对路径改成你的绝对路径(保留 <source> 那一层) | |
| python3 - <<'EOF' | |
| import json | |
| ROOT = '/your/path/to/TOSC-archive/' # 末尾带斜杠 | |
| p = 'data/qwen2-vl-2b-masked-split1-2.json' | |
| d = json.load(open(p)) | |
| for r in d: | |
| r['images'] = [ROOT + i for i in r['images']] | |
| json.dump(d, open('train.json', 'w'), ensure_ascii=False) | |
| EOF | |
| # 3) 训练脚本见 github.com/julyanghar/TOSC | |
| ``` | |
| --- | |
| ## 数据是怎么来的 | |
| masked-image 流水线(详见代码仓库 `docs/01-data-generation.md`): | |
| ``` | |
| 检测物体并返回标签 → 按标签生成遮挡图 → 挑出会引发幻觉的 → 生成句子级偏好对 → 转训练格式 | |
| ``` | |
| 源图来自 Visual Genome 与 MS-COCO train2014(公开数据集)。本归档的 `images/` 里是**遮挡处理后的图**,不是原图,所以不必另外下载源数据集就能复现训练。 | |
| --- | |
| ## 校验 | |
| ```bash | |
| # 数据 json 的图片路径是否都存在(应全部 missing=0) | |
| python3 - <<'EOF' | |
| import json, os, glob | |
| for p in sorted(glob.glob('data/*.json')): | |
| rows = json.load(open(p)); refs = set() | |
| for r in rows: | |
| for im in (r.get('images') or []): refs.add(im) | |
| for k in ('image', 'image_path', 'id'): | |
| v = r.get(k) | |
| if isinstance(v, str) and v.startswith('images/'): refs.add(v) | |
| miss = sum(1 for x in refs if not os.path.exists(x)) | |
| print(f'{os.path.basename(p):<52} {len(rows):>6} 条 {len(refs):>5} 图 missing={miss}') | |
| EOF | |
| ``` | |
| --- | |
| ## 引用 | |
| ```bibtex | |
| @article{sentinel, | |
| title = {Mitigating Object Hallucinations via Sentence-Level Early Intervention}, | |
| url = {https://github.com/pspdada/SENTINEL} | |
| } | |
| ``` | |