larpbert / README.md
abullard1's picture
Initial release of larpBERT router adapter trio (beta-binomial champion, binomial, softmax-rps)
fd9ac7d verified
|
Raw History Blame Contribute Delete
7.63 kB
metadata
base_model: answerdotai/ModernBERT-base
library_name: peft
pipeline_tag: text-classification
tags:
  - lora
  - peft
  - modernbert
  - llm-routing
  - cost-quality-pareto
  - mmlu-pro
datasets:
  - localgate/mmlu-pro-open
license: mit
language:
  - en
metrics:
  - kappa_w
  - accuracy
model-index:
  - name: larpbert
    results:
      - task:
          type: text-classification
          name: LLM Routing Difficulty Classification
        dataset:
          name: localgate/mmlu-pro-open
          type: localgate/mmlu-pro-open
        metrics:
          - name: Out-of-Fold Mean Log Loss (NLL)
            type: loss
            value: 0.5887
          - name: Linear Cohen's Kappa (kappa_w)
            type: kappa_w
            value: 0.3154
          - name: Macro F1
            type: f1
            value: 0.6957
          - name: Spearman Rank Correlation
            type: spearman
            value: 0.5001

larpBERT: Learned Adaptive Routing Predictor (ModernBERT-LoRA)

larpBERT is a parameter-efficient routing classifier for hybrid local/cloud LLM inference systems. Built by fine-tuning ModernBERT-base with low-rank adaptation (LoRA, $r=8$), larpBERT predicts query-level solvability for lightweight local models (e.g. Gemma 4 E2B) before dispatching to expensive frontier cloud endpoints.

This repository hosts the tournament champion (beta_binomial) at the root level, alongside the two runner-up finalist adapters in dedicated subfolders (binomial/ and softmax-rps/).


1. Model Summary

  • Base Architecture: answerdotai/ModernBERT-base (22 layers, 768 hidden dimension, 149M total parameters)
  • Adaptation Method: PEFT LoRA ($r = 8$, $\alpha = 16$, dropout $= 0.1$, targeting Wqkv)
  • Trainable Parameters: 1,123,590 parameters (~0.75% of base model)
  • Output Space: 6 ordinal probability bands $\ell \in {0, 1, 2, 3, 4, 5}$ corresponding to the probability that a lightweight local model solves the incoming query correctly.
  • Context Length: 1,024 tokens (native ModernBERT architecture with rotary embeddings, unpadding, and FlashAttention-2 support).

2. Tournament Finalists & Evaluation

The models were evaluated in a rigorous 108-configuration screening tournament using 5-fold cross-validation on $n = 6{,}705$ stratified development questions from localgate/mmlu-pro-open. Selection was governed by out-of-fold log loss (a proper scoring rule to guard against overconfident border pushing), with linear Cohen's $\kappa_w$ as the primary ordinal metric.

Model Variant Objective Function Best Epoch OOF Mean Log Loss $\downarrow$ Linear $\kappa_w$ $\uparrow$ Macro F1 $\uparrow$ Spearman $\rho$ $\uparrow$ HF Location
Beta-Binomial (Champion) Overdispersed Count Likelihood 8 0.58867 0.3154 0.6957 0.5001 Root (localgate/larpbert)
Binomial Standard Binomial Likelihood 5 0.59068 0.3023 0.6908 0.4962 localgate/larpbert (subfolder="binomial")
Softmax-RPS Ranked Probability Score 6 0.59160 0.3102 0.6912 0.4955 localgate/larpbert (subfolder="softmax-rps")

Note: All three finalists substantially outperform dense embedding baselines (BGE-small Centroid $\kappa_w = 0.1918$) and non-parametric Category Mean baselines ($\kappa_w = 0.1979$).


3. The 6-Level Probability Binning Scheme

Ground truth labels were constructed by evaluating Gemma 4 E2B across $k = 5$ independent stochastic passes ($T = 1.0, p = 0.95, k = 64$) graded by a 3-judge LLM panel on AWS Bedrock (2-of-3 majority consensus). The empirical success count $c \in {0, 1, 2, 3, 4, 5}$ gives continuous rate $\hat{p} = c / 5$.

The unit interval is mapped into 6 equal-width bands: β„“=min⁑(⌊6p^βŒ‹,5)∈{0,1,2,3,4,5}\ell = \min(\lfloor 6\hat{p}\rfloor, 5) \in \{0, 1, 2, 3, 4, 5\}

Level ($\ell$) Empirical Probability Band Semantic Meaning Recommended Action
0 $0.0% \le \hat{p} < 16.7%$ Extremely unlikely local success Route to Cloud Frontier
1 $16.7% \le \hat{p} < 33.3%$ Very unlikely local success Route to Cloud Frontier
2 $33.3% \le \hat{p} < 50.0%$ Unlikely local success Route to Cloud Frontier
3 $50.0% \le \hat{p} < 66.7%$ Likely local success Execute Locally
4 $66.7% \le \hat{p} < 83.3%$ Very likely local success Execute Locally
5 $83.3% \le \hat{p} \le 100.0%$ Extremely likely local success Execute Locally

4. Quickstart: Usage & Inference

Installation

pip install transformers peft torch

Loading the Champion (Root)

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from peft import PeftModel

repo_id = "localgate/larpbert"
base_model_id = "answerdotai/ModernBERT-base"

# 1. Load Tokenizer & Base ModernBERT
tokenizer = AutoTokenizer.from_pretrained(repo_id)
base_model = AutoModelForSequenceClassification.from_pretrained(
    base_model_id, 
    num_labels=6
)

# 2. Load the Champion Adapter (Beta-Binomial)
model = PeftModel.from_pretrained(base_model, repo_id)
model.eval()

# 3. Classify an incoming prompt
query = "What is the rank of a 3x3 real symmetric matrix with eigenvalues 2, 0, -1?"
inputs = tokenizer(query, return_tensors="pt", max_length=1024, truncation=True)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)
    pred_level = torch.argmax(probs, dim=-1).item()

# Standard routing threshold (Level >= 3 -> Local; Level < 3 -> Cloud)
decision = "LOCAL" if pred_level >= 3 else "CLOUD"
print(f"Predicted Difficulty Band: {pred_level} | Routing Decision: {decision}")

Loading Finalist Runner-Up Variants (subfolder)

To load the Binomial or Softmax-RPS models from the same repository:

# Load Binomial Finalist
binomial_model = PeftModel.from_pretrained(
    base_model, 
    repo_id, 
    subfolder="binomial"
)

# Load Softmax-RPS Finalist
softmax_rps_model = PeftModel.from_pretrained(
    base_model, 
    repo_id, 
    subfolder="softmax-rps"
)

5. Training Details

  • Hardware: RunPod NVIDIA A100-SXM4 (80GB VRAM)
  • Batch Size: 32 (effective batch size 32, zero gradient accumulation delay)
  • Optimizer: AdamW ($\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}$)
  • Learning Rate: $5 \times 10^{-5}$ with linear warmup across the first 10% of training steps and cosine decay
  • Weight Decay: $0.01$
  • Dropout: $0.1$ on attention and LoRA layers
  • Sequence Length: 1,024 tokens (zero prompt truncation across all training and evaluation instances)

6. Datasets & Reproducibility

  • Training Corpus: localgate/mmlu-pro-open ($n = 6{,}705$ development prompts, $5{,}848$ train / $857$ validation).
  • Code & Verification Pipeline: Full evaluation manifests, cross-validation out-of-fold predictions, noise ceiling models, and paper drafts are maintained in the LocalGate GitHub repository.

7. Citation

@article{localgate2026,
  title={Learned Adaptive Routing for Hybrid Local-Cloud LLM Inference},
  author={LocalGate Research Team},
  year={2026},
  journal={arXiv preprint},
  url={https://github.com/abullard1/localgate}
}