| --- |
| license: apache-2.0 |
| pipeline_tag: image-to-image |
| tags: |
| - image-editing |
| - multi-modal |
| --- |
| |
| # Text-Vision Co-Instructed Image Editing (TV-Edit) |
|
|
| This repository contains the weights for **TV-Edit**, a Textual-Visual instruction unified Editing framework. TV-Edit jointly models textual instructions as semantic intent and sparse visual instructions (such as points or drags) as spatial guidance to achieve precise and intent-faithful image manipulation. |
|
|
| [**Project Page**](https://xiechenxi99.github.io/TVEdit/) | [**Paper**](https://huggingface.co/papers/2606.16767) | [**GitHub**](https://github.com/PolyU-VCLab/TVEdit) |
|
|
| ## Introduction |
| TV-Edit is designed to unify the strengths of textual instructions and visual prompts. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. |
|
|
| ## Installation |
|
|
| To set up the environment, run the following commands: |
|
|
| ```shell |
| ## git clone the repository |
| git clone https://github.com/PolyU-VCLab/TVEdit.git |
| cd TVEdit |
| |
| # create an environment |
| conda create -n TVEdit python=3.10 |
| conda activate TVEdit |
| |
| pip install --upgrade pip |
| pip install torch==2.5.0+cu121 torchvision==0.20.0+cu121 --index-url https://download.pytorch.org/whl/cu121 |
| pip install transformers==4.52.4 pytorch-lightning==2.4.0 diffusers==0.35.1 |
| ``` |
|
|
| ## Usage |
|
|
| You can run the interactive Gradio demo to test the model: |
|
|
| 1. Download the base model checkpoint: [Qwen-Image-Edit](https://huggingface.co/Qwen/Qwen-Image-Edit). |
| 2. Download the trained TV-Edit weights from this repository. |
| 3. [Optional] TV-Edit supports existing trained acceleration LoRA for 4-step editing: [Qwen-Image-Edit-4step](https://huggingface.co/lightx2v/Qwen-Image-Lightning/blob/main/Qwen-Image-Edit-Lightning-4steps-V1.0-bf16.safetensors). |
| 4. Launch the Gradio demo: |
|
|
| ```shell |
| python gradio_demo.py |
| ``` |
|
|
| After launching the demo, you can upload an image, draw point trajectories for spatial control, and enter textual instructions for semantic changes. |
|
|
| ## Citation |
|
|
| If you find this work helpful, please cite: |
|
|
| ```bibtex |
| @article{xie2026text-vision, |
| title={Text-Vision Co-Instructed Image Editing}, |
| author={Xie, Chenxi and Wu, Yuhui and Yi, Qiaosi and Zhang, Lei}, |
| journal={arXiv preprint arXiv:2606.16767}, |
| year={2026}, |
| } |
| ``` |