taewhan commited on
Commit
aed26c9
·
verified ·
1 Parent(s): 0fc183e

docs(README): correct demo caption to point-tracking via feature similarity

Browse files

The GIF is a point-query tracking demo (query point relayed by argmax of last-frame cosine-similarity), not generic self-attention. Fix caption accordingly.

Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -24,7 +24,7 @@ performance across image and video benchmarks — and leads on DAVIS video track
24
  <p align="center">
25
  <img src="assets/haaland_full_attn_blk20.gif" width="480" alt="Self-attention visualization on a video clip"/>
26
  </p>
27
- <p align="center"><em>Self-attention over patch tokens on a video clip — the encoder attends coherently to the subject across frames.</em></p>
28
 
29
  - **Architecture**: ViT-7B (embed 4096 / depth 40 / heads 32), patch 16, 3D axial RoPE
30
  (`base=100`), SwiGLU FFN, LayerScale, per-head QK-norm, gated attention, 4 register tokens.
 
24
  <p align="center">
25
  <img src="assets/haaland_full_attn_blk20.gif" width="480" alt="Self-attention visualization on a video clip"/>
26
  </p>
27
+ <p align="center"><em>Point tracking on a video clip — a query point propagated across frames by patch-feature cosine similarity.</em></p>
28
 
29
  - **Architecture**: ViT-7B (embed 4096 / depth 40 / heads 32), patch 16, 3D axial RoPE
30
  (`base=100`), SwiGLU FFN, LayerScale, per-head QK-norm, gated attention, 4 register tokens.