xrag-7b-router / README.md
wexumin's picture
Update README.md
f6a4da1 verified
|
Raw
History Blame Contribute Delete
1.45 kB
---
license: apache-2.0
tags:
- context-compression
- routing
- xrag
- retrieval-augmented-generation
base_model: Hannibal046/xrag-7b
library_name: compression_router
---
# xRAG Router
A routing classifier for [xRAG](https://huggingface.co/Hannibal046/xrag-7b) that decides at inference time whether compressed context is sufficient or whether to fall back to the full context.
Trained with [overflowguard](https://github.com/s-nlp/overflowguard/tree/main) on SQuAD using an LLM judge (DeepSeek) for evaluation, with 5-fold stratified CV for threshold selection (Youden's J statistic).
## Usage
```python
from examples.train_xrag import XragRouter
# Load base model + routing classifier in one call
router = XragRouter.from_pretrained("wexumin/xrag-7b-router")
router.park_gpu()
# Auto-routing: classifier decides compressed vs full
result = router.run_pipeline("long document text...", query="what is X?")
print(result["prediction"]) # answer
print(result["mode"]) # "compressed" or "full"
print(result["tokens_saved"])
```
## What's in this repo
- `router_config.json` — base model path + routing threshold
- `routing_clf.pt` — classifier weights (2-layer MLP, ~2K params)
The base xRAG model is loaded automatically from `Hannibal046/xrag-7b`. This repo only contains the lightweight routing classifier on top.
This model also requires custom code from [xRAG](https://github.com/Hannibal046/xRAG/tree/main) repository.