| --- |
| license: apache-2.0 |
| tags: |
| - context-compression |
| - routing |
| - xrag |
| - retrieval-augmented-generation |
| base_model: Hannibal046/xrag-7b |
| library_name: compression_router |
| --- |
| |
| # xRAG Router |
|
|
| A routing classifier for [xRAG](https://huggingface.co/Hannibal046/xrag-7b) that decides at inference time whether compressed context is sufficient or whether to fall back to the full context. |
|
|
| Trained with [overflowguard](https://github.com/s-nlp/overflowguard/tree/main) on SQuAD using an LLM judge (DeepSeek) for evaluation, with 5-fold stratified CV for threshold selection (Youden's J statistic). |
|
|
| ## Usage |
|
|
| ```python |
| from examples.train_xrag import XragRouter |
| |
| # Load base model + routing classifier in one call |
| router = XragRouter.from_pretrained("wexumin/xrag-7b-router") |
| router.park_gpu() |
| |
| # Auto-routing: classifier decides compressed vs full |
| result = router.run_pipeline("long document text...", query="what is X?") |
| print(result["prediction"]) # answer |
| print(result["mode"]) # "compressed" or "full" |
| print(result["tokens_saved"]) |
| ``` |
|
|
| ## What's in this repo |
|
|
| - `router_config.json` — base model path + routing threshold |
| - `routing_clf.pt` — classifier weights (2-layer MLP, ~2K params) |
|
|
| The base xRAG model is loaded automatically from `Hannibal046/xrag-7b`. This repo only contains the lightweight routing classifier on top. |
| This model also requires custom code from [xRAG](https://github.com/Hannibal046/xRAG/tree/main) repository. |