Open-Sourced model and data for ULTRAIF: Advancing Instruction Following from the Wild.
li sheng
bambisheng
AI & ML interests
None yet
Recent Activity
upvoted a paper 2 days ago
EasyPPO: Stabilizing the Critic Is Key upvoted a paper 8 days ago
Improving Test-Time Scaling with Adaptive Looped Transformers upvoted a paper 26 days ago
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks