viavicdev's picture
Model card: redact conversion recipe, correct sizes/usage, add measured example
d2df1d1 verified
|
Raw
History Blame Contribute Delete
6.8 kB
---
license: apache-2.0
base_model: omni-research/Tarsier2-Recap-7b
base_model_relation: quantized
pipeline_tag: video-text-to-text
library_name: mlx
language:
- en
tags:
- mlx
- apple-silicon
- video-captioning
- qwen2-vl
- tarsier
- 4-bit
---
# Tarsier2-Recap-7b-MLX-4bit
A 4-bit [MLX](https://github.com/ml-explore/mlx) conversion of
[omni-research/Tarsier2-Recap-7b](https://huggingface.co/omni-research/Tarsier2-Recap-7b),
prepared for native inference on Apple Silicon.
Tarsier2-Recap-7b produces unusually specific, grounded video descriptions — it
tends to name concrete details rather than summarise a scene generically. No
official MLX build exists. This repository provides one: the same weights,
quantized to 4-bit and run on the Mac GPU through MLX.
> The underlying model is in the Qwen2-VL 7B class. Hugging Face may display a
> lower automatic parameter count (~1.9 B) because MLX 4-bit weights are packed
> into 32-bit integers.
- **Base model:** omni-research/Tarsier2-Recap-7b (Apache-2.0)
- **Format:** MLX, 4-bit affine quantization (group size 64)
- **Model weights:** approximately 5.64 GB / 5.25 GiB (two safetensors shards)
- **Complete repository:** approximately 5.65 GB / 5.26 GiB
- **Underlying architecture:** Qwen2-VL 7B (see conversion notes)
- **Not retrained or fine-tuned** — format conversion and quantization only
## Requirements
Verified with [`mlx-vlm`](https://github.com/Blaizzy/mlx-vlm) 0.6.8 on macOS
15.7.2, Python 3.12, `mlx` 0.32.0.
```bash
pip install mlx-vlm==0.6.8
brew install ffmpeg # for extracting frames
```
## Usage
```python
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("viavicdev/Tarsier2-Recap-7b-MLX-4bit")
# Provide the video as a list of extracted frame images (see note 1)
frames = ["frame_01.jpg", "frame_02.jpg", "frame_03.jpg", "frame_04.jpg"]
prompt = apply_chat_template(
processor, model.config,
"Describe this video in detail.",
num_images=len(frames), # must match len(frames)
)
result = generate(
model, processor, prompt,
image=frames, # parameter is image= (not images=)
max_tokens=300,
temperature=0.3,
repetition_penalty=1.3, # see note 3
repetition_context_size=40,
)
print(result.text) # .text — generate() returns a GenerationResult
```
Extract frames with ffmpeg, for example eight evenly spaced frames from a clip
of known duration `D`:
```bash
ffmpeg -i clip.mp4 -vf "fps=8/D,scale=-2:448" -frames:v 8 -q:v 2 frame_%02d.jpg
```
### Practical notes
1. **Pass frames as images, not a video file.** The `mlx-vlm` video decoder path
is unreliable for this model and can produce output unrelated to the clip.
Extracting frames with ffmpeg and passing them as a list (`image=[...]` with
a matching `num_images`) delivers the actual frames to the model.
2. **The chat template is included** in this repository (`chat_template.json`
and `chat_template.jinja`). Some MLX repackings omit it, which makes
`apply_chat_template` fail.
3. **Apply a repetition penalty.** At 4-bit precision the model can occasionally
loop. `repetition_penalty=1.3`, `repetition_context_size=40` and
`temperature=0.3` kept output stable in our testing; in a small internal set
of 22 short clips, one degenerated without a penalty and none did with it.
4. **Runtime scales with frame resolution and frame count**, not with the
duration of the source clip. See the measurements below.
## Example
A single measured run, not a benchmark.
- **Input:** 8 frames (796×448) from a 10.7-second clip —
[*Trains – Mini World Lyon*](https://commons.wikimedia.org/wiki/File:Trains_-_Mini_World_Lyon.webm)
by Benoît Prieur, Wikimedia Commons, CC0
- **Hardware:** Apple M4 Pro, 24 GB unified memory, macOS 15.7.2
- **Versions:** Python 3.12.12, `mlx` 0.32.0, `mlx-vlm` 0.6.8
- **Model load (cold, excluded from generation timing):** ≈14 s
- **Generation:** median **15.8 s** over 3 warm runs (15.3 / 15.8 / 16.1)
- **Peak memory:** ≈6.8 GB
- **Settings:** `max_tokens=300`, as in the usage example above
Output (run 2 of 3, verbatim):
> A model train, consisting of a green engine and several red freight cars,
> moves along the tracks from left to right. The background features a large
> crowd of people gathered in an open area with numerous colorful cars parked
> and displayed. The train passes by the car display area multiple times, moving
> smoothly along the tracks.
For comparison, 8 frames at 252×448 (a vertical clip) on the same machine ran in
a median of **5.5 s**. Frame resolution dominates runtime — budget accordingly.
## Conversion notes
The original checkpoint could not be converted directly with the standard
`mlx-vlm` workflow, because its published architecture metadata does not fully
match the underlying multimodal model structure.
For this release, the checkpoint was adapted to a compatible native MLX layout
before 4-bit quantization. This required resolving differences in the
vision-language architecture and in the positional encoding configuration.
The resulting checkpoint runs through the standard native inference path in
`mlx-vlm`. The conversion changes the storage and runtime format only; the model
was not retrained or fine-tuned.
The complete conversion procedure is not included in this repository.
## Limitations
- **4-bit quantization trades some fidelity** for reduced size and higher speed.
For maximum quality, use the original bf16 weights on a CUDA device.
- **English only.** It describes in English regardless of any spoken language in
the clip.
- **Vision only (no audio).** It captions what is seen, not what is heard.
- **Free-form description over rigid schemas.** It follows open-ended captioning
prompts more reliably than strict output formats, and is better used as a
captioner than as a classifier.
- **Detail depends on the source.** On low-resolution or distant subjects it
describes what is visible at that scale and may not identify the subject
specifically.
- **Not systematically evaluated here.** For accuracy figures, see the base
model's card and paper. The timings above are single measurements on one
machine.
## License and attribution
This conversion inherits the **Apache-2.0** license of the base model.
All credit for the original model, training and weights belongs to the Tarsier2
authors ([omni-research/Tarsier2-Recap-7b](https://huggingface.co/omni-research/Tarsier2-Recap-7b)).
This repository provides the MLX format conversion, 4-bit quantization,
packaging, Apple Silicon compatibility testing and usage documentation. The
model was not retrained or fine-tuned.
When citing Tarsier2, cite the original authors.