Papers
arxiv:2609.37402

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Published on Sep 29
· Submitted by
Lucius
on Sep 30
Authors:
,
,
,
,
,
,

Abstract

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.

Community

Paper author Paper submitter
•
edited about 7 hours ago

Hi everyone! We’re the SinapisAI team, and we’re introducing SaveRouter, our work on economical LLM routing in collaboration with researchers from Nanjing University and HKUST.

LLM routing can lower inference bills, but collecting the feedback needed to train a router costs money too. We study how many deployment queries it takes for routing savings to recover that upfront cost.

SaveRouter selectively collects informative query–model feedback and shares what it learns across related queries, with additional refinement for individual queries.

Across four routing benchmarks, our main setting:

  • Uses only ~33–41% of the available training feedback.
  • Maintains competitive or better routing quality.
  • Reduces the number of queries needed to break even by a factor of ~1.9–9.5 compared with the earliest-paying conventional fully supervised baseline, when matching the best single model’s quality.

One takeaway: the supervision budget that gives the lowest serving cost isn’t necessarily the one that pays back earliest.

Happy to discuss the method, evaluation, or how these trade-offs arise in your own routing setups!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37402
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.37402 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.37402 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.37402 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.