Papers
arxiv:2609.39071

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

Published on Sep 30
· Submitted by
YidaCai
on Oct 5
Authors:
,
,
,

Abstract

Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

Community

Paper author Paper submitter

We introduce LexReward, a taxonomy-driven framework for legal reward modeling that evaluates response quality across three complementary dimensions: Style, Element, and Chain. We develop dimension-specific rubrics and use the resulting rewards to construct pairwise preference data for DPO and to train a family of reward models, LexRM. Experiments show that the rubric-based rewards reliably distinguish response quality and that DPO improves performance across all three dimensions. LexRM further enables reinforcement learning to improve policy performance in each corresponding dimension without requiring reference answers at reward time.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.39071
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39071 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39071 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39071 in a Space README.md to link it from this page.

Collections including this paper 1