Predict human preference to LLM responses.
Binfeng Xu PRO
billxbf
AI & ML interests
evolving back to apes
Recent Activity
upvoted a paper about 18 hours ago
OSWorld-Pro: Process-based Evaluation for Computer Use Agents upvoted a paper about 19 hours ago
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches commentedon a paper 3 days ago
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General AgentsOrganizations
None yet