File size: 2,954 Bytes
844bc93
f40d97f
 
 
 
844bc93
61609da
844bc93
f40d97f
 
 
844bc93
 
f40d97f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
61801ee
 
 
 
 
 
f40d97f
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
title: Molmo2Fish Tracking
emoji: 🐟
colorFrom: purple
colorTo: blue
sdk: gradio
sdk_version: 6.15.1
app_file: app.py
short_description: Track fish in sonar video, fix it with plain English
python_version: "3.12"
startup_duration_timeout: 1h
---

# 🐟 Molmo2Fish β€” interactive fish tracking with natural language guidance

Demo of [**tidalove/Molmo2Fish**](https://huggingface.co/tidalove/Molmo2Fish), the model from
*"Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance"*
([paper](https://huggingface.co/papers/2608.18602) Β·
[code](https://github.com/tidalove/molmo2fish)).

Molmo2Fish is a LoRA-finetuned [Molmo2](https://huggingface.co/allenai) VLM that tracks
salmon in ARIS **sonar** video and β€” crucially β€” accepts a plain-English critique of its
own output and re-emits corrected tracks. The paper reports tracking accuracy going from
~5% to ~79% over a handful of conversational correction turns.

## How the demo works

1. **β‘  Track all fish** sends the clip with the prompt `track all fish`. The model replies
   with the html-v2 pointing format used in training:
   `<tracks coords="0.0 1 409 852;0.5 1 436 890 2 300 120;…">fish</tracks>`
   (timestamps in seconds at 2 FPS, ids, and x/y normalised to 0–1000).
2. **β‘‘ Apply correction** rebuilds the conversation β€” video on the first user turn, the
   model's previous `<tracks …>` answer as the assistant turn, your critique as the new
   user turn β€” exactly as in `olmo/eval/vllm_runner.py::build_multi_turn_chat`, and the
   model regenerates the track set.

Points are parsed with the same regexes as
`olmo/preprocessing/point_formatter.py` and drawn back onto the 6 FPS source video.

## Example clips

The three sonar clips shipped with this Space are re-encoded (G channel, 6 FPS, libx264,
matching the repo's `encode_frames_to_video` recipe) from the Caltech Fish Counting
release [**perona-lab/cfc26**](https://huggingface.co/datasets/perona-lab/cfc26), which is
distributed under **CC-BY-4.0** β€” credit to the Caltech Fish Counting / CFC26 authors.
The correction prompts pre-filled with each example are verbatim from the validation split
of [tidalove/cfc-track-instruction](https://huggingface.co/datasets/tidalove/cfc-track-instruction).

## Notes

- Runs on ZeroGPU; the model is loaded in bfloat16 (~16 GB VRAM).
- Video is sampled at 2 FPS, max 128 frames, per the model's
  `video_preprocessor_config.json`.
- The released checkpoint ships a config mismatch: `processor_config.json` sets
  `use_frame_special_tokens: true` (so each frame is wrapped in
  `<frame_start>`/`<frame_end>`, matching training) while `config.json` sets it
  `false`, which makes the model count `<im_end>` instead and abort with
  `AssertionError: Expected 0 videos, but got 1`. `app.py` aligns the two at
  load time.
- A truncated answer (no closing `</tracks>`) means you hit the *Max new tokens* cap β€”
  raise it in **Advanced**.