JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Paper • 2607.27670 • Published 9 days ago • 2
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 3 days ago • 320
N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Paper • 2607.23783 • Published 18 days ago • 47
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Paper • 2607.02501 • Published Jul 2 • 59
HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing Paper • 2606.13898 • Published Jun 11 • 5
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference Paper • 2605.25191 • Published May 24 • 5
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks Paper • 2605.19147 • Published May 18 • 3
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274