| --- |
| license: mit |
| pipeline_tag: video-classification |
| tags: |
| - video |
| - action-recognition |
| - efficient-inference |
| - token-selection |
| library_name: pytorch |
| --- |
| |
| # LookWhen ViT-B/16 |
|
|
| Pre-trained checkpoint for **"LookWhen? Fast Video Recognition by Learning When, Where, |
| and What to Compute"** ([arXiv:2605.06809](https://arxiv.org/abs/2605.06809)). |
|
|
| LookWhen splits video recognition in two: a shallow **selector** scores every patch of a |
| 2x-downscaled (2x in each dimension) clip, and a deep **extractor** processes only the top-K of them while still |
| predicting features of the whole video. |
|
|
| This is the model used throughout the paper: ViT-B/16, pre-trained for 20 epochs on |
| Kinetics-400 and SSv2 by distilling InternVideo2's video token and DINOv3's frame and |
| patch tokens, with selection supervised by top1-distance token uniqueness. It was trained |
| over sparsity 70-95%, so it runs anywhere in that range without retraining. |
|
|
| ## Usage |
|
|
| ```bash |
| pip install git+https://github.com/alisalamatian1/LookWhen.git |
| ``` |
|
|
| ```python |
| import torch |
| from lookwhen import LookWhen |
| from lookwhen.data.video import load_clip |
| |
| model = LookWhen.from_pretrained().cuda().eval() # downloads this checkpoint |
| clip = load_clip("video.mp4", num_frames=16, img_size=224).unsqueeze(0).cuda() |
| |
| out = model(clip, keep_ratio=0.1) # keep_ratio = 1 - sparsity |
| out["video"] # (1, 768) video embedding |
| ``` |
|
|
| Code, training, and evaluation: https://github.com/alisalamatian1/LookWhen |
|
|
| ## Citation |
|
|
| ```bibtex |
| @article{salamatian2026lookwhen, |
| title = {LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute}, |
| author = {Salamatian, Ali and Fuller, Anthony and Sarkar, Pritam and |
| Green, James R. and Sigal, Leonid and Shelhamer, Evan}, |
| journal = {arXiv preprint arXiv:2605.06809}, |
| year = {2026} |
| } |
| ``` |
|
|