RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs Paper • 2609.12552 • Published 16 days ago • 5
Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures Paper • 2608.02271 • Published Aug 3
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models Paper • 2510.13537 • Published Oct 15, 2025
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published Aug 20 • 12
Falcon Perception-HD: High Density Perception via Reinforcement Learning Paper • 2608.18881 • Published Aug 19
PixRestore: Unified Image Restoration via Pixel Diffusion Transformer Paper • 2608.16793 • Published Aug 17 • 4
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding Paper • 2608.17402 • Published Aug 18 • 20
A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples Paper • 2607.29122 • Published Jul 31 • 6
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paper • 2607.23909 • Published Jul 27 • 8
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 78
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published Jul 13 • 43
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 33
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Paper • 2607.07386 • Published Jul 8 • 12
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Paper • 2607.04884 • Published Jul 6 • 10
Unified Audio Intelligence Without Regressing on Text Intelligence Paper • 2607.05196 • Published Jul 6 • 24
Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation Paper • 2606.17188 • Published Jun 17 • 2
FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes Paper • 2606.20769 • Published Jun 18 • 1