Video-Text-to-Text
Transformers
Safetensors
qwen3_vl
image-text-to-text
camera-movement
video-understanding
qwen3-vl
vggt-injection
Instructions to use ddz16/CamInject-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ddz16/CamInject-4B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ddz16/CamInject-4B") model = AutoModelForMultimodalLM.from_pretrained("ddz16/CamInject-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
metadata
base_model: Qwen/Qwen3-VL-4B-Instruct
license: apache-2.0
library_name: transformers
pipeline_tag: video-text-to-text
tags:
- camera-movement
- video-understanding
- qwen3-vl
- sft
- vggt-injection
CamInject-4B
This repository contains the CamInject-4B model from the paper Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation.
Project page: https://ddz16.github.io/cammotion.github.io
GitHub: https://github.com/ddz16/CamDistill
Camera-movement VGGT-Direct 注入 SFT 微调模型,基于 Qwen/Qwen3-VL-4B-Instruct。
- Checkpoint:
checkpoint-1326 - 训练框架: ms-swift
使用
from transformers import AutoModelForCausalLM, AutoProcessor
model = AutoModelForCausalLM.from_pretrained("ddz16/CamInject-4B", torch_dtype="bfloat16", device_map="auto")
processor = AutoProcessor.from_pretrained("ddz16/CamInject-4B")