TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 16 days ago • 139
stabilityai/stable-video-diffusion-img2vid-xt Image-to-Video • 2B • Updated Jul 10, 2024 • 150k • 3.38k
RobotisSW/Task_900011_900012_pick_place_peanut_stage3_MCAP_merged_lerobot_v30 Updated 28 days ago • 159 • 1
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering Paper • 2607.01420 • Published Jul 1 • 13
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring Paper • 2605.30834 • Published May 29 • 11
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 241