--- base_model: Qwen/Qwen3-VL-4B-Instruct license: apache-2.0 library_name: transformers pipeline_tag: video-text-to-text tags: - camera-movement - video-understanding - qwen3-vl - distillation --- # CamDistill-4B Camera-movement understanding model trained with **Camera Token Distillation** on top of `Qwen/Qwen3-VL-4B-Instruct`. A lightweight Camera Token Module learns geometry-aware camera tokens (distilled from VGGT) and injects them into the language model. Given a video, it outputs structured JSON describing every camera-movement segment. - **Paper**: [Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation](https://huggingface.co/papers/2608.10932) - **Project page**: https://ddz16.github.io/cammotion.github.io - **Code**: https://github.com/ddz16/CamDistill > ⚠️ **This model cannot be loaded with plain 🤗 Transformers.** It contains an extra Camera Token > Module and a patched forward pass. Loading it as a standard `Qwen3VLForConditionalGeneration` > would silently drop those weights and produce incorrect results. Use the CamDistill repo, which > registers the required custom model type through a plugin. ## Usage Clone the [CamDistill repo](https://github.com/ddz16/CamDistill), then run (camera tokens are generated internally — **no online VGGT required**): ```bash python camera_movement_sft/infer_single.py \ --model ddz16/CamDistill-4B \ --video /path/to/video.mp4 \ --variant camdistill ``` See the repo's README for environment setup and batch evaluation.