Language Models Meet World Models: Embodied Experiences Enhance Language Models Paper • 2305.10626 • Published May 18, 2023 • 1
On the Feasibility of Cross-Task Transfer with Model-Based Reinforcement Learning Paper • 2210.10763 • Published Oct 19, 2022 • 1
OmniControlNet: Dual-stage Integration for Conditional Image Generation Paper • 2406.05871 • Published Jun 9, 2024
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation Paper • 2508.00728 • Published Aug 1, 2025
FrontierCS: Evolving Challenges for Evolving Intelligence Paper • 2512.15699 • Published Dec 17, 2025 • 5
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents Paper • 2601.16973 • Published 3 days ago • 21
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents Paper • 2601.16973 • Published 3 days ago • 21
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning Paper • 2511.21688 • Published Nov 26, 2025 • 8
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis Paper • 2501.04377 • Published Jan 8, 2025 • 14
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time Paper • 2408.13233 • Published Aug 23, 2024 • 23
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs Paper • 2406.18521 • Published Jun 26, 2024 • 30
TokenCompose: Grounding Diffusion with Token-level Supervision Paper • 2312.03626 • Published Dec 6, 2023 • 5
TokenCompose: Grounding Diffusion with Token-level Supervision Paper • 2312.03626 • Published Dec 6, 2023 • 5