Predict human preference to LLM responses.
Binfeng Xu
billxbf
AI & ML interests
evolving back to apes
Organizations
None yet
models 21
billxbf/qwen3.5-4b-pi-polar
4B • Updated • 5
billxbf/qwen3.5-4b-opencode-polar
4B • Updated • 5
billxbf/qwen3.5-4b-qwencode-polar
4B • Updated • 21
billxbf/qwen3.5-4b-claudecode-polar
4B • Updated • 6
billxbf/qwen3.5-4b-codex-polar-step72
Reinforcement Learning • 5B • Updated • 6
billxbf/zephyr-7b-dpo-iter1
Text Generation • 274k • Updated • 12
billxbf/zephyr-7b-dpo-iter3
Text Generation • 266k • Updated • 7
billxbf/zephyr-7b-dpo-iter2
Text Generation • 266k • Updated • 5
billxbf/Nano-Raccoon-Preview-1104
425k • Updated • 4
billxbf/zephyr-7b-sft-iter3
Text Generation • 266k • Updated • 6
datasets 20
billxbf/math_pile_v3
Viewer • Updated • 1.52M • 38
billxbf/ultrafeedback-dpo-iter3
Viewer • Updated • 20.4k • 18
billxbf/ultrafeedback-dpo-iter1
Viewer • Updated • 20.4k • 18
billxbf/ultrafeedback-dpo-iter2
Viewer • Updated • 20.4k • 14
billxbf/ultrafeedback-sft-iter3
Viewer • Updated • 20.4k • 24
billxbf/ultrafeedback-sft-iter2
Viewer • Updated • 20.4k • 23
billxbf/ultrafeedback-sft-iter1
Viewer • Updated • 20.4k • 15
billxbf/verified100-chitchat
Viewer • Updated • 100 • 18
billxbf/verified100-lite
Viewer • Updated • 100 • 23
billxbf/verified100
Viewer • Updated • 100 • 34