Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
Abstract
In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.
Community
Rerankers, reward models and multi-document QA systems can score several candidates together in one LLM prompt. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, what a reader answers, or which response is preferred.
This paper shows that equal ranking quality does not imply equal decisions. On passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when the candidates are reordered. No prompt-time change tested resolves this.
The paper introduces order-consistency SFT (OC-SFT), which reduces this order dependence in the weights by penalizing disagreement between a candidate's scores across orderings. It:
• holds ranking quality
• leads every decision-stability measure among trained scorers, across passage reranking, response ranking and multi-document QA
• is more stable than order-averaged distillation on 12 base models across three families
• with a single permutation, retains sets that overlap more than ten averaged permutations of an off-the-shelf scorer
💡 Takeaway: a comparison of such scorers should report what a threshold retains and a reader answers, not ranking quality alone.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages (2026)
- ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker (2026)
- AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research (2026)
- Regime Boundary Alignment for Evidence-Gated Question Answering (2026)
- Return or Revise? Learning When Revision Helps Retrieval-Augmented QA (2026)
- Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict (2026)
- Can a Cacheable Decision Model Follow Rules? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.26762 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper