loom-video / README.md
multimodalart's picture
multimodalart HF Staff
Upload folder using huggingface_hub
ac69be1 verified
|
Raw
History Blame Contribute Delete
970 Bytes
---
title: LoomVideo
emoji: 🎬
colorFrom: green
colorTo: pink
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
short_description: Unified multimodal video generation and editing (5B)
python_version: "3.12"
startup_duration_timeout: 1h
---
# LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing
LoomVideo is a compact 5B-parameter unified architecture for both video generation and editing, built on MLLM + DiT. It supports four tasks within a single model:
- **Text-to-Video (t2v)**: Generate a video from a text prompt
- **Instruction Editing (edit)**: Edit a video following text instructions
- **Instruction-Image Editing (ref_edit)**: Edit a video with a reference image
- **Multi-Image-to-Video (mi2v)**: Compose multiple reference images into a video
## Paper
[LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing](https://arxiv.org/abs/2606.06042)
## Model
[MSALab/LoomVideo](https://huggingface.co/MSALab/LoomVideo)