Instructions to use localgate/larpbert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use localgate/larpbert with PEFT:
from peft import PeftModel from transformers import AutoModelForSequenceClassification base_model = AutoModelForSequenceClassification.from_pretrained("answerdotai/ModernBERT-base") model = PeftModel.from_pretrained(base_model, "localgate/larpbert") - Notebooks
- Google Colab
- Kaggle
Download README.md from localgate/larpbert: direct link, hf CLI and curl.
- Browser
- Download file 7.63 kB
-
https://huggingface.co/localgate/larpbert/resolve/main/README.md
- Command line
-
hf download hf://localgate/larpbert/README.md
-
curl -L -o README.md https://huggingface.co/localgate/larpbert/resolve/main/README.md
base_model: answerdotai/ModernBERT-base
library_name: peft
pipeline_tag: text-classification
tags:
- lora
- peft
- modernbert
- llm-routing
- cost-quality-pareto
- mmlu-pro
datasets:
- localgate/mmlu-pro-open
license: mit
language:
- en
metrics:
- kappa_w
- accuracy
model-index:
- name: larpbert
results:
- task:
type: text-classification
name: LLM Routing Difficulty Classification
dataset:
name: localgate/mmlu-pro-open
type: localgate/mmlu-pro-open
metrics:
- name: Out-of-Fold Mean Log Loss (NLL)
type: loss
value: 0.5887
- name: Linear Cohen's Kappa (kappa_w)
type: kappa_w
value: 0.3154
- name: Macro F1
type: f1
value: 0.6957
- name: Spearman Rank Correlation
type: spearman
value: 0.5001
larpBERT: Learned Adaptive Routing Predictor (ModernBERT-LoRA)
larpBERT is a parameter-efficient routing classifier for hybrid local/cloud LLM inference systems. Built by fine-tuning ModernBERT-base with low-rank adaptation (LoRA, $r=8$), larpBERT predicts query-level solvability for lightweight local models (e.g. Gemma 4 E2B) before dispatching to expensive frontier cloud endpoints.
This repository hosts the tournament champion (beta_binomial) at the root level, alongside the two runner-up finalist adapters in dedicated subfolders (binomial/ and softmax-rps/).
1. Model Summary
- Base Architecture:
answerdotai/ModernBERT-base(22 layers, 768 hidden dimension, 149M total parameters) - Adaptation Method: PEFT LoRA ($r = 8$, $\alpha = 16$, dropout $= 0.1$, targeting
Wqkv) - Trainable Parameters: 1,123,590 parameters (~0.75% of base model)
- Output Space: 6 ordinal probability bands $\ell \in {0, 1, 2, 3, 4, 5}$ corresponding to the probability that a lightweight local model solves the incoming query correctly.
- Context Length: 1,024 tokens (native ModernBERT architecture with rotary embeddings, unpadding, and FlashAttention-2 support).
2. Tournament Finalists & Evaluation
The models were evaluated in a rigorous 108-configuration screening tournament using 5-fold cross-validation on $n = 6{,}705$ stratified development questions from localgate/mmlu-pro-open. Selection was governed by out-of-fold log loss (a proper scoring rule to guard against overconfident border pushing), with linear Cohen's $\kappa_w$ as the primary ordinal metric.
| Model Variant | Objective Function | Best Epoch | OOF Mean Log Loss $\downarrow$ | Linear $\kappa_w$ $\uparrow$ | Macro F1 $\uparrow$ | Spearman $\rho$ $\uparrow$ | HF Location |
|---|---|---|---|---|---|---|---|
| Beta-Binomial (Champion) | Overdispersed Count Likelihood | 8 | 0.58867 | 0.3154 | 0.6957 | 0.5001 | Root (localgate/larpbert) |
| Binomial | Standard Binomial Likelihood | 5 | 0.59068 | 0.3023 | 0.6908 | 0.4962 | localgate/larpbert (subfolder="binomial") |
| Softmax-RPS | Ranked Probability Score | 6 | 0.59160 | 0.3102 | 0.6912 | 0.4955 | localgate/larpbert (subfolder="softmax-rps") |
Note: All three finalists substantially outperform dense embedding baselines (BGE-small Centroid $\kappa_w = 0.1918$) and non-parametric Category Mean baselines ($\kappa_w = 0.1979$).
3. The 6-Level Probability Binning Scheme
Ground truth labels were constructed by evaluating Gemma 4 E2B across $k = 5$ independent stochastic passes ($T = 1.0, p = 0.95, k = 64$) graded by a 3-judge LLM panel on AWS Bedrock (2-of-3 majority consensus). The empirical success count $c \in {0, 1, 2, 3, 4, 5}$ gives continuous rate $\hat{p} = c / 5$.
The unit interval is mapped into 6 equal-width bands:
| Level ($\ell$) | Empirical Probability Band | Semantic Meaning | Recommended Action |
|---|---|---|---|
| 0 | $0.0% \le \hat{p} < 16.7%$ | Extremely unlikely local success | Route to Cloud Frontier |
| 1 | $16.7% \le \hat{p} < 33.3%$ | Very unlikely local success | Route to Cloud Frontier |
| 2 | $33.3% \le \hat{p} < 50.0%$ | Unlikely local success | Route to Cloud Frontier |
| 3 | $50.0% \le \hat{p} < 66.7%$ | Likely local success | Execute Locally |
| 4 | $66.7% \le \hat{p} < 83.3%$ | Very likely local success | Execute Locally |
| 5 | $83.3% \le \hat{p} \le 100.0%$ | Extremely likely local success | Execute Locally |
4. Quickstart: Usage & Inference
Installation
pip install transformers peft torch
Loading the Champion (Root)
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from peft import PeftModel
repo_id = "localgate/larpbert"
base_model_id = "answerdotai/ModernBERT-base"
# 1. Load Tokenizer & Base ModernBERT
tokenizer = AutoTokenizer.from_pretrained(repo_id)
base_model = AutoModelForSequenceClassification.from_pretrained(
base_model_id,
num_labels=6
)
# 2. Load the Champion Adapter (Beta-Binomial)
model = PeftModel.from_pretrained(base_model, repo_id)
model.eval()
# 3. Classify an incoming prompt
query = "What is the rank of a 3x3 real symmetric matrix with eigenvalues 2, 0, -1?"
inputs = tokenizer(query, return_tensors="pt", max_length=1024, truncation=True)
with torch.no_grad():
logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
pred_level = torch.argmax(probs, dim=-1).item()
# Standard routing threshold (Level >= 3 -> Local; Level < 3 -> Cloud)
decision = "LOCAL" if pred_level >= 3 else "CLOUD"
print(f"Predicted Difficulty Band: {pred_level} | Routing Decision: {decision}")
Loading Finalist Runner-Up Variants (subfolder)
To load the Binomial or Softmax-RPS models from the same repository:
# Load Binomial Finalist
binomial_model = PeftModel.from_pretrained(
base_model,
repo_id,
subfolder="binomial"
)
# Load Softmax-RPS Finalist
softmax_rps_model = PeftModel.from_pretrained(
base_model,
repo_id,
subfolder="softmax-rps"
)
5. Training Details
- Hardware: RunPod NVIDIA A100-SXM4 (80GB VRAM)
- Batch Size: 32 (effective batch size 32, zero gradient accumulation delay)
- Optimizer: AdamW ($\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}$)
- Learning Rate: $5 \times 10^{-5}$ with linear warmup across the first 10% of training steps and cosine decay
- Weight Decay: $0.01$
- Dropout: $0.1$ on attention and LoRA layers
- Sequence Length: 1,024 tokens (zero prompt truncation across all training and evaluation instances)
6. Datasets & Reproducibility
- Training Corpus: localgate/mmlu-pro-open ($n = 6{,}705$ development prompts, $5{,}848$ train / $857$ validation).
- Code & Verification Pipeline: Full evaluation manifests, cross-validation out-of-fold predictions, noise ceiling models, and paper drafts are maintained in the LocalGate GitHub repository.
7. Citation
@article{localgate2026,
title={Learned Adaptive Routing for Hybrid Local-Cloud LLM Inference},
author={LocalGate Research Team},
year={2026},
journal={arXiv preprint},
url={https://github.com/abullard1/localgate}
}