EVO-WAM: Evolving World Action Models through Video-Action Verification Paper • 2609.38057 • Published 3 days ago • 29
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Paper • 2406.18629 • Published Jun 26, 2024 • 42