Video-Text-to-Text
MLX
Safetensors
English
qwen2_vl
apple-silicon
video-captioning
qwen2-vl
tarsier
4-bit precision
Instructions to use viavicdev/Tarsier2-Recap-7b-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use viavicdev/Tarsier2-Recap-7b-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Tarsier2-Recap-7b-MLX-4bit viavicdev/Tarsier2-Recap-7b-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
| license: apache-2.0 | |
| base_model: omni-research/Tarsier2-Recap-7b | |
| base_model_relation: quantized | |
| pipeline_tag: video-text-to-text | |
| library_name: mlx | |
| language: | |
| - en | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - video-captioning | |
| - qwen2-vl | |
| - tarsier | |
| - 4-bit | |
| # Tarsier2-Recap-7b-MLX-4bit | |
| A 4-bit [MLX](https://github.com/ml-explore/mlx) conversion of | |
| [omni-research/Tarsier2-Recap-7b](https://huggingface.co/omni-research/Tarsier2-Recap-7b), | |
| prepared for native inference on Apple Silicon. | |
| Tarsier2-Recap-7b produces unusually specific, grounded video descriptions — it | |
| tends to name concrete details rather than summarise a scene generically. No | |
| official MLX build exists. This repository provides one: the same weights, | |
| quantized to 4-bit and run on the Mac GPU through MLX. | |
| > The underlying model is in the Qwen2-VL 7B class. Hugging Face may display a | |
| > lower automatic parameter count (~1.9 B) because MLX 4-bit weights are packed | |
| > into 32-bit integers. | |
| - **Base model:** omni-research/Tarsier2-Recap-7b (Apache-2.0) | |
| - **Format:** MLX, 4-bit affine quantization (group size 64) | |
| - **Model weights:** approximately 5.64 GB / 5.25 GiB (two safetensors shards) | |
| - **Complete repository:** approximately 5.65 GB / 5.26 GiB | |
| - **Underlying architecture:** Qwen2-VL 7B (see conversion notes) | |
| - **Not retrained or fine-tuned** — format conversion and quantization only | |
| ## Requirements | |
| Verified with [`mlx-vlm`](https://github.com/Blaizzy/mlx-vlm) 0.6.8 on macOS | |
| 15.7.2, Python 3.12, `mlx` 0.32.0. | |
| ```bash | |
| pip install mlx-vlm==0.6.8 | |
| brew install ffmpeg # for extracting frames | |
| ``` | |
| ## Usage | |
| ```python | |
| from mlx_vlm import load, generate | |
| from mlx_vlm.prompt_utils import apply_chat_template | |
| model, processor = load("viavicdev/Tarsier2-Recap-7b-MLX-4bit") | |
| # Provide the video as a list of extracted frame images (see note 1) | |
| frames = ["frame_01.jpg", "frame_02.jpg", "frame_03.jpg", "frame_04.jpg"] | |
| prompt = apply_chat_template( | |
| processor, model.config, | |
| "Describe this video in detail.", | |
| num_images=len(frames), # must match len(frames) | |
| ) | |
| result = generate( | |
| model, processor, prompt, | |
| image=frames, # parameter is image= (not images=) | |
| max_tokens=300, | |
| temperature=0.3, | |
| repetition_penalty=1.3, # see note 3 | |
| repetition_context_size=40, | |
| ) | |
| print(result.text) # .text — generate() returns a GenerationResult | |
| ``` | |
| Extract frames with ffmpeg, for example eight evenly spaced frames from a clip | |
| of known duration `D`: | |
| ```bash | |
| ffmpeg -i clip.mp4 -vf "fps=8/D,scale=-2:448" -frames:v 8 -q:v 2 frame_%02d.jpg | |
| ``` | |
| ### Practical notes | |
| 1. **Pass frames as images, not a video file.** The `mlx-vlm` video decoder path | |
| is unreliable for this model and can produce output unrelated to the clip. | |
| Extracting frames with ffmpeg and passing them as a list (`image=[...]` with | |
| a matching `num_images`) delivers the actual frames to the model. | |
| 2. **The chat template is included** in this repository (`chat_template.json` | |
| and `chat_template.jinja`). Some MLX repackings omit it, which makes | |
| `apply_chat_template` fail. | |
| 3. **Apply a repetition penalty.** At 4-bit precision the model can occasionally | |
| loop. `repetition_penalty=1.3`, `repetition_context_size=40` and | |
| `temperature=0.3` kept output stable in our testing; in a small internal set | |
| of 22 short clips, one degenerated without a penalty and none did with it. | |
| 4. **Runtime scales with frame resolution and frame count**, not with the | |
| duration of the source clip. See the measurements below. | |
| ## Example | |
| A single measured run, not a benchmark. | |
| - **Input:** 8 frames (796×448) from a 10.7-second clip — | |
| [*Trains – Mini World Lyon*](https://commons.wikimedia.org/wiki/File:Trains_-_Mini_World_Lyon.webm) | |
| by Benoît Prieur, Wikimedia Commons, CC0 | |
| - **Hardware:** Apple M4 Pro, 24 GB unified memory, macOS 15.7.2 | |
| - **Versions:** Python 3.12.12, `mlx` 0.32.0, `mlx-vlm` 0.6.8 | |
| - **Model load (cold, excluded from generation timing):** ≈14 s | |
| - **Generation:** median **15.8 s** over 3 warm runs (15.3 / 15.8 / 16.1) | |
| - **Peak memory:** ≈6.8 GB | |
| - **Settings:** `max_tokens=300`, as in the usage example above | |
| Output (run 2 of 3, verbatim): | |
| > A model train, consisting of a green engine and several red freight cars, | |
| > moves along the tracks from left to right. The background features a large | |
| > crowd of people gathered in an open area with numerous colorful cars parked | |
| > and displayed. The train passes by the car display area multiple times, moving | |
| > smoothly along the tracks. | |
| For comparison, 8 frames at 252×448 (a vertical clip) on the same machine ran in | |
| a median of **5.5 s**. Frame resolution dominates runtime — budget accordingly. | |
| ## Conversion notes | |
| The original checkpoint could not be converted directly with the standard | |
| `mlx-vlm` workflow, because its published architecture metadata does not fully | |
| match the underlying multimodal model structure. | |
| For this release, the checkpoint was adapted to a compatible native MLX layout | |
| before 4-bit quantization. This required resolving differences in the | |
| vision-language architecture and in the positional encoding configuration. | |
| The resulting checkpoint runs through the standard native inference path in | |
| `mlx-vlm`. The conversion changes the storage and runtime format only; the model | |
| was not retrained or fine-tuned. | |
| The complete conversion procedure is not included in this repository. | |
| ## Limitations | |
| - **4-bit quantization trades some fidelity** for reduced size and higher speed. | |
| For maximum quality, use the original bf16 weights on a CUDA device. | |
| - **English only.** It describes in English regardless of any spoken language in | |
| the clip. | |
| - **Vision only (no audio).** It captions what is seen, not what is heard. | |
| - **Free-form description over rigid schemas.** It follows open-ended captioning | |
| prompts more reliably than strict output formats, and is better used as a | |
| captioner than as a classifier. | |
| - **Detail depends on the source.** On low-resolution or distant subjects it | |
| describes what is visible at that scale and may not identify the subject | |
| specifically. | |
| - **Not systematically evaluated here.** For accuracy figures, see the base | |
| model's card and paper. The timings above are single measurements on one | |
| machine. | |
| ## License and attribution | |
| This conversion inherits the **Apache-2.0** license of the base model. | |
| All credit for the original model, training and weights belongs to the Tarsier2 | |
| authors ([omni-research/Tarsier2-Recap-7b](https://huggingface.co/omni-research/Tarsier2-Recap-7b)). | |
| This repository provides the MLX format conversion, 4-bit quantization, | |
| packaging, Apple Silicon compatibility testing and usage documentation. The | |
| model was not retrained or fine-tuned. | |
| When citing Tarsier2, cite the original authors. | |