human preference dataset stanfordnlp/SHP-2 Viewer • Updated Jan 11, 2024 • 4.07M • 1.01k • 17 Anthropic/hh-rlhf Viewer • Updated May 26, 2023 • 169k • 34.4k • 1.7k OpenMOSS-Team/hh-rlhf-strength-cleaned Viewer • Updated Jan 31, 2024 • 168k • 122 • 23 heegyu/hh-rlhf-vicuna-format Viewer • Updated Sep 6, 2023 • 169k • 3 • 4
RM OpenAssistant/reward-model-deberta-v3-large-v2 Text Classification • Updated Feb 1, 2023 • 14.9k • • 244
OpenAssistant/reward-model-deberta-v3-large-v2 Text Classification • Updated Feb 1, 2023 • 14.9k • • 244
human preference dataset stanfordnlp/SHP-2 Viewer • Updated Jan 11, 2024 • 4.07M • 1.01k • 17 Anthropic/hh-rlhf Viewer • Updated May 26, 2023 • 169k • 34.4k • 1.7k OpenMOSS-Team/hh-rlhf-strength-cleaned Viewer • Updated Jan 31, 2024 • 168k • 122 • 23 heegyu/hh-rlhf-vicuna-format Viewer • Updated Sep 6, 2023 • 169k • 3 • 4
RM OpenAssistant/reward-model-deberta-v3-large-v2 Text Classification • Updated Feb 1, 2023 • 14.9k • • 244
OpenAssistant/reward-model-deberta-v3-large-v2 Text Classification • Updated Feb 1, 2023 • 14.9k • • 244