# TOSC — Triadic Object-State Consistency Benchmark, evaluation code, and trained LoRA models for the **Triadic Object-State Consistency (TOSC)** benchmark. TOSC tests whether a vision-language model keeps its object descriptions consistent across three versions of the same scene: the **original** image, the same image with one object **removed** (inpainted out), and the same image with that object **replaced** by a different one. A consistent model should mention the object when present, stop mentioning it once removed, and report the new object after replacement. This repository is **self-contained**: the benchmark images, the evaluation code, and the model adapters are all here, and the evaluation pipeline runs with no external code dependency. ## Repository layout ``` tosc-model-weights/ ├── benchmark/ # 5,969 triplets (origin / removed / replaced) │ ├── original/images/ # {N}.jpg │ ├── repair/images/ # {N}_masked_{obj}_repair.png (removed) │ ├── insert/images/ # {N}_masked_{obj}_insert_{new}.png (replaced) │ ├── TOSC_dataset.jsonl # generated index (regenerable) │ └── insertions.jsonl # generated object metadata (regenerable) ├── eval/ # evaluation code (see eval/README.md for details) │ ├── build_index.py # image folders -> the two .jsonl files │ ├── run_eval.sh # end-to-end runner: caption -> merge -> score │ ├── eval_tosc.py # scoring engine │ └── eval_masked_obj_generative.py ├── llava/ # vendored inference package (imports as `llava`) ├── lora/ # 4 trained LoRA adapters (TSA-DPO) └── requirements.txt ``` ### Models | Adapter (`lora/…`) | Base model | |--------------------|-----------| | `LLaVA_v1_5_7b-TSA-DPO` | `liuhaotian/llava-v1.5-7b` | | `LLaVA_v1_5_13b-TSA-DPO` | `liuhaotian/llava-v1.5-13b` | | `Qwen2_VL_7B-TSA-DPO` | `Qwen/Qwen2-VL-7B-Instruct` | | `Qwen2_5_VL_7B-TSA-DPO` | `Qwen/Qwen2.5-VL-7B-Instruct` | ## Installation A CUDA GPU is required for caption generation. Python 3.10+ is recommended. ```bash # 1. Get the repo (images + adapters are stored with Git LFS) git lfs install git clone https://huggingface.co/julyanghar/tosc-model-weights cd tosc-model-weights # 2. Create an environment and install dependencies conda create -n tosc python=3.10 -y && conda activate tosc # or use a venv pip install -r requirements.txt # 3. (Qwen models only) install flash-attn for the Qwen2-VL / Qwen2.5-VL path, # or change attn_implementation to "sdpa" in llava/eval/utils/hf_utils.py pip install flash-attn --no-build-isolation ``` Base models (`liuhaotian/llava-v1.5-*`, `Qwen/Qwen2*-VL-*-Instruct`) and the CLIP vision tower are downloaded automatically from the Hugging Face Hub on first use; the LoRA adapters are loaded from `lora/` in this repo. ## How to run ```bash # 1. Build the index files once (regenerable from the images at any time) python eval/build_index.py # 2. Evaluate a model. Pick one of the four presets in eval/run_eval.sh # (uncomment it), or select it via environment variables. Example: Qwen2.5-VL # on 4 GPUs. GPU_LIST=0,1,2,3 \ OUTPUT_NAME=Qwen2_5_VL_7B-TSA-DPO \ MODEL_NAME=Qwen/Qwen2.5-VL-7B-Instruct \ LORA_NAME=$PWD/lora/Qwen2_5_VL_7B-TSA-DPO \ bash eval/run_eval.sh ``` `run_eval.sh` runs three steps: (1) generate a caption for every benchmark image (multi-GPU, sharded by `GPU_LIST`), (2) merge the shards, (3) score with `eval/eval_tosc.py`. The inference path is selected automatically — model names containing `llava` + `1.5` use the LLaVA-1.5 loader, everything else uses the Qwen path. Set `LORA_NAME=""` to evaluate a base model without the adapter. ### Single-GPU / quick smoke test ```bash GPU_LIST=0 bash eval/run_eval.sh # single GPU (slower) ``` To sanity-check the wiring on a few images, point `QUESTION_FILE` at a trimmed copy of `benchmark/TOSC_dataset.jsonl`. ### Outputs Results go to `${TOSC_RESULTS:-/home/yilin/tmp/tosc_results}` (outside the repo by default; override with `OUTPUT_DIR`): - `answers/.jsonl` — merged captions - `_eval.json` — full per-triplet results - `_eval_summary.json` — aggregate metrics ### Configuration (environment variables) | Variable | Default | Meaning | |----------|---------|---------| | `GPU_LIST` | `0` | comma-separated GPU ids, e.g. `0,1,2,3` | | `MODEL_NAME` / `LORA_NAME` / `OUTPUT_NAME` | Qwen2.5 preset | model selection | | `MAX_NEW_TOKENS` | `512` | caption length | | `TEMPERATURE` | `0` | `0` = greedy / deterministic | | `SEED` | `42` | random seed | | `OUTPUT_DIR` | `/home/yilin/tmp/tosc_results` | where results are written | ## Metrics The scorer reports the headline **TOSC** score plus **OPA** (sees the object), **RCA** (drops it after removal), **RUA** (sees the replacement, not the old object), and fine-grained failure rates (OldPersist / RepFail / CC / MixConf). Object mentions are detected with a local COCO synonym table and lemmatization — no external API or LLM judge. See [eval/README.md](eval/README.md) for the exact formulas and per-state definitions.