Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
HuggingFaceH4 's Collections
— Journal Club 📚 —
Scaling Test-Time Compute with Open Models
Zephyr ORPO
Zephyr 7B
Zephyr 7B Gemma
StarChat2 15B
Papers We've Read
Awesome SFT datasets
Awesome feedback datasets
Awesome reward models

Awesome reward models

updated 12 days ago

A curated collection of reward models to use with techniques like rejection sampling and RLHF / RLAIF

Upvote
11

  • llm-blender/PairRM

    Text Generation • 0.4B • Updated Jan 22, 2024 • 677 • 209

  • openbmb/UltraRM-13b

    Updated Oct 14, 2023 • 1.28k • 61

  • OpenAssistant/reward-model-deberta-v3-large-v2

    Text Classification • Updated Feb 1, 2023 • 13.3k • • 247

  • PKU-Alignment/beaver-7b-v1.0-reward

    Reinforcement Learning • 7B • Updated Apr 20, 2024 • 2.62k • 17
Upvote
11
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs