Succinct Router 14M

A 14M-parameter decoder-only routing model trained from random initialization. It predicts an independent pass probability for each of these candidate configurations:

  • gpt-5.6-luna-none
  • gpt-5.6-terra-low
  • gpt-5.6-sol-medium

This model does not answer prompts. It selects the cheapest candidate configuration expected to pass a calibrated quality threshold.

Evaluation

{
  "exact_route_accuracy": 0.9508599508599509,
  "unsafe_downroute_rate": 0.005528255528255528,
  "selected_model_pass_rate": 0.9944717444717445,
  "abstention_rate": 0.0,
  "savings_vs_always_large": 0.8486732186732187
}

Run with MLX

pip install "mlx>=0.32,<0.33" tokenizers
python route_mlx.py --model-dir . --prompt "Extract the invoice number from INV-4821."

The result contains calibrated pass probabilities and the selected candidate configuration. If no candidate clears the threshold, abstained_to_largest is true and the route falls back to the largest configuration.

Limitations

Training and evaluation data are synthetic and mechanically graded. The artifact is a learning prototype, not a production-ready router. Validate it on anonymized real traffic before making product or cost claims. MLX runtime parity and latency must be measured on Apple Silicon.

Downloads last month
19
Safetensors
Model size
13.8M params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using pmarquees/succinct-router 1