| viewer: false | |
| Unison is a comprehensive benchmark comprising 2,169 high-quality unified task samples, designed to evaluate joint understanding and generation in unified multimodal models. | |
| ## Tasks | |
| | Task | Name | What it measures | | |
| |------|------|------------------| | |
| | **IC** | Internal Consistency | Internal alignment between understanding and generation. | | |
| | **UGG** | Understanding Guided Generation | Model's ability to leverage comprehension guiding contextually generation. | | |
| | **GGU** | Generation Guided Understanding | How synthesized outputs assist understanding tasks. | | |
| | **ME** | Mutual Enhancement | Multi-turn iterative synergy where understanding identifies generation errors and vice versa, enabling mutual refinement. | | |
| <div style="display: flex; align-items: center; justify-content: space-between; width: 100%; gap: 2%;"> | |
| <img src="statistics.png" alt="Dataset statistics" style="width: 34%; margin: 0;" /> | |
| <img src="Bench_Compare.png" alt="Benchmark comparison" style="width: 65%; margin: 0;" /> | |
| </div> | |
| ## Layout | |
| ``` | |
| ├── Internal_Consistency/ # IC | |
| │ ├── prompts.txt # 556 generation prompts | |
| │ ├── questions.json # yes/no VQA keyed by prompt_<N> | |
| │ └── images/ # 556 reference images | |
| ├── Und_Guided_Gen/ # UGG | |
| │ ├── UGG.json # 543 grounded generation items | |
| │ └── images/ # 539 source images | |
| ├── Gen_Guided_Und/ # GGU (three spatial sub-categories) | |
| │ ├── 2D_Spatial/ # 182 2D spatial items | |
| │ ├── 3D_Spatial/ # 180 3D spatial items | |
| │ └── Complex_Relation/ # 180 relation items | |
| └── Mutual_Enhancement/ # ME | |
| ├── ME.json # 264 rows, 528 samples | |
| └── images/ # 264 source images | |
| ``` | |
| ## Field notes | |
| - **IC** — `questions.json` is keyed `prompt_<N>` matching the 0-indexed line `N` in `prompts.txt`; each value maps question numbers to yes/no questions about the generated image. | |
| - **UGG** — `image_path` is relative to `Und_Guided_Gen/` (e.g. `images/1.jpg`); `bbox`/`mask` give the grounded edit region, `operation` the edit type. | |
| - **GGU** — `image_path` is relative to each sub-category dir: `2D_Spatial/matrices/...`, `3D_Spatial/cubes/...`. `Complex_Relation` items carry no source image and are reconstructed from their `description` / validation questions. | |
| - **ME** — `image_path` is a bare filename resolving under `Mutual_Enhancement/images/`. | |
Xet Storage Details
- Size:
- 2.63 kB
- Xet hash:
- d01019a4e2d40ce076aefef52583a1b3be8dd7de4b8ba3648bb212a7a5a7d864
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.