--- base_model: answerdotai/ModernBERT-base library_name: peft pipeline_tag: text-classification tags: - lora - peft - modernbert - llm-routing - cost-quality-pareto - mmlu-pro datasets: - localgate/mmlu-pro-open license: mit language: - en metrics: - kappa_w - accuracy model-index: - name: larpbert results: - task: type: text-classification name: LLM Routing Difficulty Classification dataset: name: localgate/mmlu-pro-open type: localgate/mmlu-pro-open metrics: - name: Out-of-Fold Mean Log Loss (NLL) type: loss value: 0.5887 - name: Linear Cohen's Kappa (kappa_w) type: kappa_w value: 0.3154 - name: Macro F1 type: f1 value: 0.6957 - name: Spearman Rank Correlation type: spearman value: 0.5001 --- # larpBERT: Learned Adaptive Routing Predictor (ModernBERT-LoRA) **larpBERT** is a parameter-efficient routing classifier for hybrid local/cloud LLM inference systems. Built by fine-tuning [ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) with low-rank adaptation (LoRA, $r=8$), larpBERT predicts query-level solvability for lightweight local models (e.g. Gemma 4 E2B) before dispatching to expensive frontier cloud endpoints. This repository hosts the **tournament champion** (`beta_binomial`) at the root level, alongside the two runner-up finalist adapters in dedicated subfolders (`binomial/` and `softmax-rps/`). --- ## 1. Model Summary - **Base Architecture**: [`answerdotai/ModernBERT-base`](https://huggingface.co/answerdotai/ModernBERT-base) (22 layers, 768 hidden dimension, 149M total parameters) - **Adaptation Method**: PEFT LoRA ($r = 8$, $\alpha = 16$, dropout $= 0.1$, targeting `Wqkv`) - **Trainable Parameters**: 1,123,590 parameters (~0.75% of base model) - **Output Space**: 6 ordinal probability bands $\ell \in \{0, 1, 2, 3, 4, 5\}$ corresponding to the probability that a lightweight local model solves the incoming query correctly. - **Context Length**: 1,024 tokens (native ModernBERT architecture with rotary embeddings, unpadding, and FlashAttention-2 support). --- ## 2. Tournament Finalists & Evaluation The models were evaluated in a rigorous 108-configuration screening tournament using 5-fold cross-validation on $n = 6{,}705$ stratified development questions from [localgate/mmlu-pro-open](https://huggingface.co/datasets/localgate/mmlu-pro-open). Selection was governed by out-of-fold log loss (a proper scoring rule to guard against overconfident border pushing), with linear Cohen's $\kappa_w$ as the primary ordinal metric. | Model Variant | Objective Function | Best Epoch | OOF Mean Log Loss $\downarrow$ | Linear $\kappa_w$ $\uparrow$ | Macro F1 $\uparrow$ | Spearman $\rho$ $\uparrow$ | HF Location | |---|---|:---:|:---:|:---:|:---:|:---:|---| | **Beta-Binomial (Champion)** | Overdispersed Count Likelihood | 8 | **0.58867** | **0.3154** | **0.6957** | **0.5001** | *Root* (`localgate/larpbert`) | | **Binomial** | Standard Binomial Likelihood | 5 | 0.59068 | 0.3023 | 0.6908 | 0.4962 | `localgate/larpbert` (`subfolder="binomial"`) | | **Softmax-RPS** | Ranked Probability Score | 6 | 0.59160 | 0.3102 | 0.6912 | 0.4955 | `localgate/larpbert` (`subfolder="softmax-rps"`) | *Note: All three finalists substantially outperform dense embedding baselines (BGE-small Centroid $\kappa_w = 0.1918$) and non-parametric Category Mean baselines ($\kappa_w = 0.1979$).* --- ## 3. The 6-Level Probability Binning Scheme Ground truth labels were constructed by evaluating Gemma 4 E2B across $k = 5$ independent stochastic passes ($T = 1.0, p = 0.95, k = 64$) graded by a 3-judge LLM panel on AWS Bedrock (2-of-3 majority consensus). The empirical success count $c \in \{0, 1, 2, 3, 4, 5\}$ gives continuous rate $\hat{p} = c / 5$. The unit interval is mapped into 6 equal-width bands: $$\ell = \min(\lfloor 6\hat{p}\rfloor, 5) \in \{0, 1, 2, 3, 4, 5\}$$ | Level ($\ell$) | Empirical Probability Band | Semantic Meaning | Recommended Action | |:---:|:---:|---|---| | **0** | $0.0\% \le \hat{p} < 16.7\%$ | Extremely unlikely local success | **Route to Cloud Frontier** | | **1** | $16.7\% \le \hat{p} < 33.3\%$ | Very unlikely local success | **Route to Cloud Frontier** | | **2** | $33.3\% \le \hat{p} < 50.0\%$ | Unlikely local success | **Route to Cloud Frontier** | | **3** | $50.0\% \le \hat{p} < 66.7\%$ | Likely local success | **Execute Locally** | | **4** | $66.7\% \le \hat{p} < 83.3\%$ | Very likely local success | **Execute Locally** | | **5** | $83.3\% \le \hat{p} \le 100.0\%$ | Extremely likely local success | **Execute Locally** | --- ## 4. Quickstart: Usage & Inference ### Installation ```bash pip install transformers peft torch ``` ### Loading the Champion (Root) ```python import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer from peft import PeftModel repo_id = "localgate/larpbert" base_model_id = "answerdotai/ModernBERT-base" # 1. Load Tokenizer & Base ModernBERT tokenizer = AutoTokenizer.from_pretrained(repo_id) base_model = AutoModelForSequenceClassification.from_pretrained( base_model_id, num_labels=6 ) # 2. Load the Champion Adapter (Beta-Binomial) model = PeftModel.from_pretrained(base_model, repo_id) model.eval() # 3. Classify an incoming prompt query = "What is the rank of a 3x3 real symmetric matrix with eigenvalues 2, 0, -1?" inputs = tokenizer(query, return_tensors="pt", max_length=1024, truncation=True) with torch.no_grad(): logits = model(**inputs).logits probs = torch.softmax(logits, dim=-1) pred_level = torch.argmax(probs, dim=-1).item() # Standard routing threshold (Level >= 3 -> Local; Level < 3 -> Cloud) decision = "LOCAL" if pred_level >= 3 else "CLOUD" print(f"Predicted Difficulty Band: {pred_level} | Routing Decision: {decision}") ``` ### Loading Finalist Runner-Up Variants (`subfolder`) To load the **Binomial** or **Softmax-RPS** models from the same repository: ```python # Load Binomial Finalist binomial_model = PeftModel.from_pretrained( base_model, repo_id, subfolder="binomial" ) # Load Softmax-RPS Finalist softmax_rps_model = PeftModel.from_pretrained( base_model, repo_id, subfolder="softmax-rps" ) ``` --- ## 5. Training Details - **Hardware**: RunPod NVIDIA A100-SXM4 (80GB VRAM) - **Batch Size**: 32 (effective batch size 32, zero gradient accumulation delay) - **Optimizer**: AdamW ($\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}$) - **Learning Rate**: $5 \times 10^{-5}$ with linear warmup across the first 10% of training steps and cosine decay - **Weight Decay**: $0.01$ - **Dropout**: $0.1$ on attention and LoRA layers - **Sequence Length**: 1,024 tokens (zero prompt truncation across all training and evaluation instances) --- ## 6. Datasets & Reproducibility - **Training Corpus**: [localgate/mmlu-pro-open](https://huggingface.co/datasets/localgate/mmlu-pro-open) ($n = 6{,}705$ development prompts, $5{,}848$ train / $857$ validation). - **Code & Verification Pipeline**: Full evaluation manifests, cross-validation out-of-fold predictions, noise ceiling models, and paper drafts are maintained in the [LocalGate GitHub repository](https://github.com/abullard1/localgate). --- ## 7. Citation ```bibtex @article{localgate2026, title={Learned Adaptive Routing for Hybrid Local-Cloud LLM Inference}, author={LocalGate Research Team}, year={2026}, journal={arXiv preprint}, url={https://github.com/abullard1/localgate} } ```