metadata
license: apache-2.0
tags:
- molmo2
- codec
- adacodec
- video
- stage2
extra_gated_prompt: >-
Contains weights derived from an unreleased base (Molmo2-4B-SFT). Research
use.
molmo2-codec Stage-2 V4 (trained codec video model)
Stage-2 SFT checkpoints for the AdaCodec-on-Molmo2 video pipeline, V4 (trainable E_P P-frame ViT branch). Video is ingested as a codec representation — I-frames: 81 tokens, P-frames: 16 tokens via the trained E_P — at ~33% of dense's visual-token budget (3x compression).
Results — MLVU multiple-choice (n=100, official dense-vs-codec eval, same 100 examples)
| checkpoint | codec acc | token budget |
|---|---|---|
| v2b (frozen P-tokenizer, 0.49M) | 35.0% | ~25% |
| stage2_v4_step500.pt | 46.0% | 33.2% |
| stage2_v4_step1200.pt | 52.0% | 33.2% |
| dense (per-frame RGB, reference) | 71.0% | 100% |
The trainable ~381M E_P branch (vs the old frozen 0.49M P-tokenizer) is the key change: codec accuracy went 35% -> 46% (500 steps) -> 52% (1200 steps) and was still climbing.
Files
stage2_v4_step500.pt,stage2_v4_step1200.pt— consolidated{model, proj}state dicts (~19GB fp32, full 4.85B model incl. the trained E_P). Load withcodec/codec_eval.py --ckpt <file>(E_P auto-attaches because the ckpt carriescodec_ptok.vit_p.*).
Code: https://github.com/weikaih04/molmo2-codec (branch adacodec). Stage-1 E_P:
https://huggingface.co/weikaih/molmo2-codec-stage1-v4