RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Paper • 2607.28509 • Published • 30
None defined yet.
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
Light-WAM: Efficient World Action Models with State-Fusion Action Decoding