A comprehensive framework designed to cultivate VLMs with human-like visuospatial abilities.
Ray Yang
rayruiyang
AI & ML interests
None yet
Recent Activity
upvoted a paper 16 days ago
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone upvoted a paper about 1 month ago
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation updated a collection 2 months ago
VSTOrganizations
None yet