| --- |
| license: mit |
| library_name: DETECTRON2 |
| tags: |
| - computer-vision |
| - image-segmentation |
| - video-instance-segmentation |
| - deep-learning |
| - query-enrichment |
| - pytorch |
| datasets: |
| - youtube-vis |
| - ovis |
| pipeline_tag: image-segmentation |
| metrics: |
| - ap |
| - ar |
| base_model: |
| - deahuang/MinVIS-Model-Zoo |
| --- |
| |
| # QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment |
|
|
| This repository hosts the official model weights and evaluation benchmarks for **QueenVIS**. |
|
|
| ## π Quick Links |
|
|
| * **Project Page:** [aut-aisl.github.io/QueenVIS-Project-Page](https://aut-aisl.github.io/QueenVIS-Project-Page/) |
| * **GitHub Repository:** [github.com/ArianKheir/QueenVIS](https://github.com/ArianKheir/QueenVIS) |
| * **Paper (arXiv):** [10.48550/arXiv.2607.24598](https://doi.org/10.48550/arXiv.2607.24598) |
|
|
| --- |
|
|
| ## π Model Description |
|
|
| **QueenVIS** rethinks the paradigm of image-only training for Video Instance Segmentation (VIS). By introducing training-only auxiliary heads for feature prediction and center prediction, QueenVIS embeds dense appearance and spatial priors directly into transformer object queries without requiring video clip supervision. |
|
|
| * **Key Feature:** Adds **zero parameters** and **zero extra GFLOPs** at inference relative to image-only baselines like MinVIS. |
| * **Temporal Association:** Employs a training-free confidence-guided query propagation mechanism paired with a non-parametric memory bank to maintain temporal identity across frames. |
|
|
| --- |
|
|
| ## π Datasets & Evaluation Metrics |
|
|
| QueenVIS is evaluated across standard VIS benchmarks: |
|
|
| * **YouTube-VIS (2019 / 2021 / 2022 Long Split):** Standard and long-sequence video benchmarks. |
| * **OVIS (Occluded Video Instance Segmentation):** A challenging benchmark characterized by severe occlusion and complex object interactions. |
|
|
| Performance is measured using standard Video Instance Segmentation metrics: |
| * **AP (Average Precision):** Primary metric across IoU thresholds. |
| * **APβ
β & APββ
:** Average Precision at IoU thresholds of 0.50 and 0.75. |
| * **ARβ & ARββ:** Average Recall given 1 and 10 detections per video. |
|
|
| --- |
|
|
| ## π Evaluation Results |
|
|
| ### 1. Main Benchmarks (ResNet-50 Backbone) |
|
|
| | Dataset | AP | APβ
β | APββ
| ARβ | ARββ | |
| | :--- | :---: | :---: | :---: | :---: | :---: | |
| | **YouTube-VIS 2019** | **51.8** | 75.2 | 58.3 | 49.4 | 62.1 | |
| | **YouTube-VIS 2021** | **50.9** | 72.1 | 55.8 | 42.8 | 57.3 | |
| | **YouTube-VIS 2022 (Long)** | **33.6** | 56.6 | 34.0 | 31.7 | 39.9 | |
| | **OVIS (Occluded VIS)** | **29.8** | 52.7 | 29.0 | 15.4 | 34.6 | |
|
|
| --- |
|
|
| ### 2. Scaling to Vision Transformers (Swin-L Backbone) |
|
|
| | Dataset | AP | APβ
β | APββ
| ARβ | ARββ | |
| | :--- | :---: | :---: | :---: | :---: | :---: | |
| | **YouTube-VIS 2019** | **63.2** | 85.3 | 68.3 | 55.5 | 68.4 | |
| | **YouTube-VIS 2021** | **59.8** | 81.5 | 65.2 | 48.3 | 64.2 | |
| | **OVIS (Occluded VIS)** | **41.0** | 62.2 | 43.0 | 18.4 | 45.0 | |
|
|
| --- |
|
|
| ## π How to Use |
|
|
| For complete instructions on installation, environment setup, dataset preparation, and inference scripts, visit the official [GitHub Repository](https://github.com/ArianKheir/QueenVIS). |
|
|
| --- |
|
|
| ## π Citation |
|
|
| If you find QueenVIS useful in your research, please cite our work: |
|
|
| ```bibtex |
| @misc{kheirandish2026queenvis, |
| title={QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment}, |
| author={Arian Kheirandish and Fardin Ayar and Ehsan Javanmardi and Manabu Tsukada and Mahdi Javanmardi}, |
| year={2026}, |
| eprint={2607.24598}, |
| archivePrefix={arXiv}, |
| primaryClass={cs.CV}, |
| url={[https://doi.org/10.48550/arXiv.2607.24598](https://doi.org/10.48550/arXiv.2607.24598)} |
| } |