Collections
Discover the best community collections!
Collections trending this week
-
In deep reinforcement learning, a pruned network is a good network
Paper • 2402.12479 • Published • 19 -
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Paper • 2403.03950 • Published • 15 -
RLHF Workflow: From Reward Modeling to Online RLHF
Paper • 2405.07863 • Published • 71 -
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Paper • 2405.11143 • Published • 41
-
Aspik101/vicuna-13b-v1.5-PL-lora_unload
Text Generation • Updated • 51 • • 2 -
eryk-mazus/polka-1.1b-chat
Text Generation • 1B • Updated • 1.61k • 19 -
speakleash/Bielik-7B-Instruct-v0.1
Text Generation • 7B • Updated • 2.82k • • 64 -
speakleash/Bielik-11B-v2.2-Instruct
Text Generation • 11B • Updated • 63
-
In deep reinforcement learning, a pruned network is a good network
Paper • 2402.12479 • Published • 19 -
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Paper • 2403.03950 • Published • 15 -
RLHF Workflow: From Reward Modeling to Online RLHF
Paper • 2405.07863 • Published • 71 -
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Paper • 2405.11143 • Published • 41
-
Aspik101/vicuna-13b-v1.5-PL-lora_unload
Text Generation • Updated • 51 • • 2 -
eryk-mazus/polka-1.1b-chat
Text Generation • 1B • Updated • 1.61k • 19 -
speakleash/Bielik-7B-Instruct-v0.1
Text Generation • 7B • Updated • 2.82k • • 64 -
speakleash/Bielik-11B-v2.2-Instruct
Text Generation • 11B • Updated • 63