Papers
arxiv:2609.16816

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

Published on Sep 15
ยท Submitted by
Bowen Qin
on Sep 16
Authors:
,
,

Abstract

Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls. Rather than providing fixed rubrics, ImpossibleRubrics provides task environments and certificates, allowing rubrics to be generated downstream and then adversarially tested for whether they reward certificate-violating answers. Eleven generators are exploited 8--26% of the time on the unbiased 150-of-169 environment cut; on a deliberately selected stress cut the strongest generator we measured is still exploited 36% while a certificate-faithful rubric is exploited 0%, so what we measure is a rubric-quality gap, not task impossibility. One result runs against intuition. A single generic rubric ("be decisive, penalize hedging") used unchanged for every task is exploited 64% of the time, and seven of the eleven generators are exploited more often than that while writing a rubric tailored to each one. The tailored criteria appear to tell an attacker which claim to fabricate. The problem is not that rubrics are vague; it is that they are specific about the wrong things.

Community

Paper author Paper submitter

Evaluating Large Language Models (LLMs) via a single aggregate scalar score often masks subtle failures and leaves room for benchmark gaming. Impossible Rubrics introduces an open-source evaluation framework and dataset designed to move beyond coarse summary metrics through fine-grained, multi-dimensional, and adversarial evaluation rubrics.

By targeting edge cases that trip up state-of-the-art models in complex reasoning, strict constraint-following, and fine-grained logical judgment, the benchmark exposes hidden failure modes and reduces LLM-as-a-Judge evaluation bias.

Key Highlights:

Multi-Dimensional Criteria: Evaluates model performance across fine-grained operational rubrics rather than relying on a single overall score.

Adversarial & Hard Test Cases: Focuses on complex, edge-case scenarios designed to reveal the true limits of frontier models.

Reliable Evaluation Protocol: Implements rigorous scoring mechanisms to mitigate LLM judge biases and inconsistencies.

Fully Open-Source: Includes the dataset, evaluation pipeline, and baseline results for full community reproducibility.

๐Ÿ“Œ Project Page: https://impossiblerubrics.github.io/

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.16816
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.16816 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.16816 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.16816 in a Space README.md to link it from this page.

Collections including this paper 1