SimpleMemVLA
Collection
A Simple but Effective Native-Video Memory for Vision-Language-Action Models • 11 items • Updated • 1
This repository contains the model presented in SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models.
SimpleMemVLA is a vision-language-action (VLA) model for long-horizon robot manipulation. It keeps sampled observation history intact and feeds it to the VLM backbone as timestamped native video, eliminating dedicated memory modules while maintaining low latency via shared-prefix prefilling.