lmms-lab-encoder/LLaVA-OneVision-2-8B-Instruct Image-Text-to-Text • 9B • Updated 23 days ago • 9.7k • 14
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Paper • 2608.08469 • Published Aug 9 • 2
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 28 days ago • 189
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published Aug 12 • 43