Abstract
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.
Community
Hi everyone,
TL;DR: Rubric RL rewards usually sum the points of satisfied criteria, so different verdict patterns can get the same reward. RRT uses item response theory instead and infers each rollout's quality from its verdict pattern.
- Reward = inferred quality: a 2-parameter IRT model gives the posterior-mode quality of each rollout, so rollouts with equal point totals can still get different rewards.
- Response Parameter Network: predicts each criterion's difficulty and discrimination from the prompt and criterion text, and is updated with online EM as the policy changes.
- Adaptive Fisher selection: judges only the most informative criteria, saving judge requests.
Results (Qwen3.5-4B):
- +1.7 macro criterion score over GRPO across Medical, Science, RaR Science, and RubricBench
- +2.8 to +5.6 on medium, hard, and very hard criteria
- At half the criterion budget, within 0.1 points of GRPO with full judging, with 49% fewer judge requests
Takeaway: a rubric criterion's verdict tells you more when you know how hard and how discriminative that criterion is.
We'd love to hear your thoughts, especially on rubrics whose criteria don't all measure one shared quality.
Get this paper in your agent:
hf papers read 2609.35646 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper