Video-Text-to-Text
Transformers
Safetensors
qwen3_vl
image-text-to-text
camera-movement
video-understanding
qwen3-vl
distillation
Instructions to use ddz16/CamDistill-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ddz16/CamDistill-4B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ddz16/CamDistill-4B") model = AutoModelForMultimodalLM.from_pretrained("ddz16/CamDistill-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| base_model: Qwen/Qwen3-VL-4B-Instruct | |
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: video-text-to-text | |
| tags: | |
| - camera-movement | |
| - video-understanding | |
| - qwen3-vl | |
| - distillation | |
| # CamDistill-4B | |
| Camera-movement understanding model trained with **Camera Token Distillation** on top of | |
| `Qwen/Qwen3-VL-4B-Instruct`. A lightweight Camera Token Module learns geometry-aware camera | |
| tokens (distilled from VGGT) and injects them into the language model. Given a video, it outputs | |
| structured JSON describing every camera-movement segment. | |
| - **Paper**: [Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation](https://huggingface.co/papers/2608.10932) | |
| - **Project page**: https://ddz16.github.io/cammotion.github.io | |
| - **Code**: https://github.com/ddz16/CamDistill | |
| > ⚠️ **This model cannot be loaded with plain 🤗 Transformers.** It contains an extra Camera Token | |
| > Module and a patched forward pass. Loading it as a standard `Qwen3VLForConditionalGeneration` | |
| > would silently drop those weights and produce incorrect results. Use the CamDistill repo, which | |
| > registers the required custom model type through a plugin. | |
| ## Usage | |
| Clone the [CamDistill repo](https://github.com/ddz16/CamDistill), then run (camera tokens are generated internally — **no online | |
| VGGT required**): | |
| ```bash | |
| python camera_movement_sft/infer_single.py \ | |
| --model ddz16/CamDistill-4B \ | |
| --video /path/to/video.mp4 \ | |
| --variant camdistill | |
| ``` | |
| See the repo's README for environment setup and batch evaluation. | |