pisco-7b-router / README.md
wexumin's picture
Update README.md
e8ba7c5 verified
|
Raw
History Blame Contribute Delete
1.35 kB
---
license: apache-2.0
tags:
- context-compression
- routing
- pisco
base_model: naver/pisco-mistral
library_name: compression_router
---
# PISCO Router
A routing classifier for [PISCO](https://huggingface.co/naver/pisco-mistral/tree/main) (COCOM architecture) that decides at inference time whether compressed context is sufficient or whether to fall back to the full context.
Trained with [overflowguard](https://github.com/s-nlp/overflowguard/tree/main) on SQuAD using an LLM judge (DeepSeek) for evaluation, with 5-fold stratified CV for threshold selection (Youden's J statistic).
## Usage
```python
from examples.train_pisco import PiscoRouter
# Load base model + routing classifier in one call
router = PiscoRouter.from_pretrained("wexumin/pisco-7b-router")
router.park_gpu()
# Auto-routing: classifier decides compressed vs full
result = router.run_pipeline("long document text...", query="what is X?")
print(result["prediction"]) # answer
print(result["mode"]) # "compressed" or "full"
print(result["tokens_saved"])
```
## What's in this repo
- `router_config.json` — base model path + routing threshold
- `routing_clf.pt` — classifier weights (2-layer MLP, ~2K params)
The base PISCO model is loaded automatically from `naver/pisco-mistral`. This repo only contains the lightweight routing classifier on top.