relay-laya

Trained Laya classifier used by the laya/auto router of Relay, a coding agent harness. It answers 18 typed questions about a coding request: task type, complexity, scope, risk, ambiguity, reasoning requirement, recommended capability tier and effort, agent role, validation, required tools and security sensitivity.

Current checkpoint: v3

Published on 2026-10-08. This checkpoint continues v2, which continued v1; it is not a new model trained from scratch. The original v1 revision remains available under the v1 tag and its immutable commit. The intermediate v2 remains a local checkpoint and was not published to this repository.

  • Base: multilingual Laya using jhu-clsp/mmBERT-base.
  • Loader: laya 0.3.27.
  • Capability labels: fast, balanced, strong, frontier.
  • Incremental training: 3 epochs; 34 focused exercises (10 corrected earlier exercises and 24 provider-neutral examples), mixed with 150 replay exercises. Validation: 93 rows. Test: 99 rows, including 8 new held-out provider-neutral examples.
  • Provider-neutral examples cover Claude, OpenAI Codex, local LM Studio, GLM, Kimi and Qwen contexts. They label task capability and reasoning effort, not concrete model IDs.
  • No private project source, session transcripts, telemetry or training dataset is published here.

Routing contract and limitations

The classifier recommends task requirements. Actual model selection belongs to Relay's deterministic router and configured model catalog. The v3 weights alone do not enforce provider boundaries, authenticate providers, switch the Codex desktop model selector or make a weak model stronger.

The accompanying Relay implementation uses laya.followProvider: true and optional laya.modelGroups to stay inside the selected provider/family, including retries and direct calls. Groups distinguish families sharing the same gateway. This requires that implementation; published Relay packages containing only the original router do not gain these controls by downloading this checkpoint.

Map real available model IDs to capability tiers. Four distinct models are not required. Unmapped providers retain the selected model with unclassified capability. A frontier requirement without a sufficient model remains a limitation, not an automatic relabeling of a smaller model. Reasoning controls must match the selected model's supported capabilities.

Evaluation

These are synthetic/template-based tests, not a real-world coding benchmark. Aggregate scores count labeled answers, not successful software tasks. Some focused examples have only two labels.

Evaluation v3 Comparison
Combined held-out labeled answers 1614/1654 (97.58%) v2: 1604/1654 (96.98%)
Original v1 test, all 18 questions 1603/1638 (97.86%) v1: 1613/1638 (98.47%)
Original test capability tier 91/91 v1: 91/91
New provider-neutral held-out capability tier 7/8 v1: 6/8
New provider-neutral held-out reasoning effort 4/8 v1: 4/8

Original-test regression is 0.61 percentage point, within the project's 1 percentage-point acceptance threshold. Retention is not perfect. The new eight-example holdout is small and reasoning-effort results show room for improvement. High synthetic accuracy overstates likely real-world accuracy.

Distribution

Relay releases pin the Hugging Face commit and SHA-256 hashes in their model manifest. Updating this repository does not automatically update already-released packages or images. No npm release or container image release accompanies this model publication.

Use revision="v1" for the original checkpoint and revision="v3" for this checkpoint. Immutable commit IDs are preferred for reproducible deployments.

Downloads last month
31
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support