WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 7 days ago • 29
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation Paper • 2606.02479 • Published Jun 1 • 23
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 3 days ago • 86
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting Paper • 2603.21304 • Published Mar 22 • 33
Open-RS Collection Model weights & datasets in the paper "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t" • 8 items • Updated Mar 21, 2025 • 13
view article Article Fine-tuning SmolLM with Group Relative Policy Optimization (GRPO) by following the Methodologies prithivMLmods • Feb 17, 2025 • 30
LLaMo: Large Language Model-based Molecular Graph Assistant Paper • 2411.00871 • Published Oct 31, 2024 • 22