A research collaboration between SinapisAI and LAMDA.
Guannan Lai · Gelin Bian · Hao-Xuan Ma · Jun-Peng Jiang · Long Chen · Jian-Dong Liu · Zhi-Hao Tan · Han-Jia Ye
Nanjing University · The Hong Kong University of Science and Technology · SinapisAI
LLM routing can reduce serving cost, but training a router often requires executing many candidate models on historical queries first. SAVERouter treats this supervision expenditure as part of the routing problem. It adaptively acquires a small, informative subset of query-model feedback, shares capability information across related queries, and retains query-level corrections for fine-grained routing.
Across four routing benchmarks, SAVERouter uses roughly 33–41% of the available training feedback while preserving competitive routing quality. The paper introduces two metrics that account for both the upfront investment and the subsequent serving-time savings:
- SA-BEP: the number of deployment queries required to recover supervision expenditure.
- SA-CR: the serving-cost ratio after amortizing supervision expenditure over a fixed deployment horizon.
Overview of SAVERouter.
SAVERouter has three main stages:
- Adaptive acquisition selects exactly
Kcandidate models per training query using grouped empirical-Bayes UCB. - Hierarchical capability estimation combines a structured group-model prior with shrinkage toward the feedback collected for each group and model.
- Query-level refinement learns contextual residuals and constructs the quality-cost routing frontier.
The sparse supervision artifact contains only acquired (query, model, quality, cost) tuples. Router fitting never receives a dense outcome matrix.
Clone the repository and create the environment:
git clone https://github.com/LAMDA-Model-Reuse/SaveRouter.git
cd SaveRouter
bash scripts/setup.shRun the download-free smoke test and unit tests:
.venv/bin/saverouter smoke-test
.venv/bin/python -m pytestRun one benchmark or the complete four-benchmark suite:
bash scripts/reproduce.sh llmrouterbench
bash scripts/reproduce.sh allThe reproduction script downloads the original benchmark data and frozen
encoders, uses the paper profiles under configs/paper/, and writes results to
outputs/. A fresh full run requires approximately 12–15 GB of free disk
space. CUDA is recommended for MMR-Bench and first-time feature extraction;
CPU execution is supported.
After installation, the equivalent CLI command is:
saverouter reproduce --benchmark all --device auto --verifyThe paper's main comparison is shown below. Small numerical differences can
occur across BLAS, CUDA, and encoder environments; --verify checks all
metrics under the repository's declared tolerances.
Main results on four routing benchmarks.
The evaluator constructs the complete policy family used in the paper: 201 cost-weight policies, cost-threshold policies, and incremental predicted-quality-gain per predicted-cost gates.
outputs/<benchmark>/result.json # profile, metrics, and metadata
outputs/<benchmark>/pareto.csv # physical routing frontier
outputs/<benchmark>/supervision.npz # exact sparse observations
outputs/main_results.csv # four-benchmark summary
Dataset revisions, model order, query order, split, seed, grouping profile, and acquisition mask are recorded for reproducibility. Dataset and encoder artifacts are downloaded from their publishers at pinned revisions.
from saverouter import FixedKSparseRouter, simulate_fixed_k_supervision
feedback = simulate_fixed_k_supervision(
rewards_train,
group_ids,
costs=costs_train,
k=4,
seed=42,
)
router = FixedKSparseRouter().fit(query_features, feedback)
choices = router.route(
test_features,
max_cost=0.01,
group_ids=test_group_ids,
)For live feedback acquisition, use collect_fixed_k_supervision with a callback
that invokes a model only after it is selected. See
examples/online_feedback.py.
Let C0 be upfront supervision expenditure, Cb the serving cost of the best
single model, Cr routed serving cost at the quality target, and Co online
routing overhead:
SA-BEP = ceil(C0 / (Cb - Cr - Co))
SA-CR@H = (C0 + H * (Cr + Co)) / (H * Cb)
SA-CR@H < 1 indicates that routing has paid back its upfront supervision cost
by deployment horizon H. The experiments use H = 1,000,000.
configs/paper/ frozen benchmark profiles
examples/ custom and online-feedback examples
results/reference/ numerical reproduction references
saverouter/ method, benchmark adapters, and CLI
scripts/ setup and reproduction entry points
tests/ unit and leakage-regression tests
If you find SAVERouter useful, please cite:
@article{lai2026routing,
title = {Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing},
author = {Lai, Guannan and Bian, Gelin and Ma, Hao-Xuan and Jiang, Jun-Peng and Chen, Long and Liu, Jian-Dong and Tan, Zhi-Hao and Ye, Han-Jia},
journal = {arXiv preprint arXiv:2609.37402},
year = {2026},
url = {https://arxiv.org/abs/2609.37402}
}This repository adapts benchmark loaders and evaluation conventions from ORBIT. Benchmark datasets and pretrained encoders retain their respective licenses and terms.
SAVERouter is released under the MIT License.

