--- license: apache-2.0 tags: - context-compression - routing - pisco base_model: naver/pisco-mistral library_name: compression_router --- # PISCO Router A routing classifier for [PISCO](https://huggingface.co/naver/pisco-mistral/tree/main) (COCOM architecture) that decides at inference time whether compressed context is sufficient or whether to fall back to the full context. Trained with [overflowguard](https://github.com/s-nlp/overflowguard/tree/main) on SQuAD using an LLM judge (DeepSeek) for evaluation, with 5-fold stratified CV for threshold selection (Youden's J statistic). ## Usage ```python from examples.train_pisco import PiscoRouter # Load base model + routing classifier in one call router = PiscoRouter.from_pretrained("wexumin/pisco-7b-router") router.park_gpu() # Auto-routing: classifier decides compressed vs full result = router.run_pipeline("long document text...", query="what is X?") print(result["prediction"]) # answer print(result["mode"]) # "compressed" or "full" print(result["tokens_saved"]) ``` ## What's in this repo - `router_config.json` — base model path + routing threshold - `routing_clf.pt` — classifier weights (2-layer MLP, ~2K params) The base PISCO model is loaded automatically from `naver/pisco-mistral`. This repo only contains the lightweight routing classifier on top.