Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published Jul 7 • 31
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 42
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 18
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 18
BRAID Collection Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process • 1 item • Updated Jul 7 • 1
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 18
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 42
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 42
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe Paper • 2605.03677 • Published May 5 • 28
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning Paper • 2605.06326 • Published May 7 • 26