Spaces:
Running on Zero
Running on Zero
| title: TimeChat-Captioner | |
| emoji: 🎬 | |
| colorFrom: purple | |
| colorTo: yellow | |
| sdk: gradio | |
| sdk_version: 6.19.0 | |
| app_file: app.py | |
| short_description: Time-aware audio-visual captions for multi-scene videos | |
| python_version: "3.12" | |
| startup_duration_timeout: 1h | |
| # TimeChat-Captioner | |
| Generate detailed, time-aware, and structurally coherent audio-visual captions for | |
| multi-scene videos. Upload a short video clip (~60 seconds recommended) and get a | |
| script-like description with timestamps that coordinates visual and audio information. | |
| ## Model | |
| [yaolily/TimeChat-Captioner-GRPO-7B](https://huggingface.co/yaolily/TimeChat-Captioner-GRPO-7B) — | |
| a 7B VLM fine-tuned from Qwen2.5-Omni-7B for dense video captioning. | |
| ## Paper | |
| [TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions](https://arxiv.org/abs/2602.08711) (ICML 2026) | |
| ## Usage | |
| 1. Upload a video (recommended ~60 second clips for best timestamps). | |
| 2. Optionally customize the prompt. | |
| 3. Click "Generate Caption" to get a structured, time-aware description. |