laya-conductor
A fine-tuned laya checkpoint for coding agent routing decisions β trained on real sessions from the omp coding agent.
Drop into the omp-conductor plugin or call directly via the laya-mlx API.
What it does
Instead of paying a frontier LLM to decide "is this request trivial or complex?", this model answers in ~21ms locally with no tokens spent.
Three decisions per call, returned as calibrated probabilities:
| Question | Type | Options |
|---|---|---|
role |
choice |
smol Β· default Β· slow |
should_continue |
noul |
P(agent stopped mid-task) |
route_tier maps directly to thinking levels in omp:
smolβ:off(no thinking, fast)defaultβ baselineslowβ:max(deep reasoning)
Performance
Fine-tuned on 4,105 real coding agent turns (1,648 route + 2,457 continuation), labeled by Claude Sonnet. Trained 4 epochs on a T4 GPU (~30 min).
| laya base (zero-shot) | laya-conductor | |
|---|---|---|
| Routing accuracy | ~37% | 87% |
| Latency (hot, M4 Pro) | 35β320 ms | 21 ms |
| Tokens spent | 0 | 0 |
Loss curve: 0.65 β 0.57 β 0.50 β 0.45 across 4 epochs.
Calibration temperatures: choice=1.361 Β· score=1.2 Β· noul=1.467
Live routing examples
21ms smol p=0.80 rename variable x to y
20ms smol p=0.78 write a git commit message
20ms default p=0.61 add OAuth login with Google
20ms default p=0.62 write unit test for parseDate
21ms slow p=0.85 why does the server crash under load? investigate
20ms slow p=0.53 design a new multi-tenant billing system
Quick start
With laya-mlx (recommended β no server needed)
pip install git+https://github.com/mizorewww/laya-mlx.git
import laya_mlx as laya
agent = laya.load("mvilacad/laya-conductor", dtype="float16")
result = agent.predict(
{"request": "why does the auth service randomly crash under load?"},
{
"role": {
"type": "choice",
"instructions": "How much reasoning effort does this coding request need?",
"criteria": {
"smol": "trivial mechanical edit: rename, typo, format, commit message, one-liner",
"default": "normal coding: implement a feature, fix a known bug, write a test",
"slow": "hard: debug unknown cause, architecture, design, large refactor",
},
}
},
)
print(result["answers"]["role"]["choice"]) # β slow
print(result["answers"]["role"]["probabilities"]) # β {"smol": 0.05, "default": 0.10, "slow": 0.85}
With laya-serve (HTTP, compatible with Jev SDK)
uv tool install 'laya[serve]'
LAYA_MODELS=conductor LAYA_PORT=48384 laya-serve
from typesafe_sdk import Choice, TypeSafeClient
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:48384")
r = client.system_one(
state={"request": "add dark mode toggle to the settings page"},
questions={"role": Choice(
instructions="How much reasoning effort does this coding request need?",
criteria={"smol": "trivial/mechanical", "default": "normal coding", "slow": "hard/unknown"},
)},
)
print(r.choices["role"].choice) # β default
Questions & state format
The model expects state as a JSON object with a request field (the user's message, max ~400 chars). Additional fields like context_tokens or objective improve accuracy when available.
{
"state": {
"request": "refactor the auth module to support multiple providers",
"context_tokens": 28000
}
}
For should_continue, pass the tail of the last assistant response:
{
"state": {
"last_response": "I've updated auth.ts. Next I'll update the tests and...",
"task": "refactor auth module"
}
}
Training
Fine-tuned from convaiinnovations/laya using RLCD (the upstream training recipe) on real omp coding agent sessions:
- Dataset: 4,105 items from 2,141 real sessions (private, contains proprietary code)
- Labels: generated by Claude Sonnet (
kr/claude-sonnet-4.6) as teacher with soft target distributions - Hardware: T4 GPU via Google Colab CLI, ~30 min
- Recipe: 4 epochs, lr encoder=2.5e-5 / head=1e-4, batch=64, GRPO + soft CE, cosine schedule
- Notebook:
conductor_finetune_colab.ipynbin this repo β plug in your own sessions and retrain
The dataset is private (contains real coding sessions), but the training notebook is fully reproducible with your own data. Collect sessions from your own coding agent, label with any LLM, and fine-tune with the notebook.
Model details
| Base model | convaiinnovations/laya (ModernBERT-large, 421M) |
| Checkpoint | mvilacad/laya-conductor |
| Size | ~843 MB |
| Context | 1,024 tokens |
| Precision | fp16 |
| Device | CPU / MPS (Apple Silicon) / CUDA |
Calibration temperatures were fitted on a held-out slice (400 items) after training. The temperature_by_options bucket from the base checkpoint is removed β it would silently mask the new calibration.
Limitations
- Zero-shot accuracy on domains far from coding (e.g. customer support, legal) will be lower than on coding tasks.
- "Slow" is over-represented in the training data (~66%) because real sessions skew toward complex tasks. If your workflow has more trivial tasks, a short fine-tune on your own data will help.
context_tokensas a downgrade guard (never switch to a lower tier when context is large) is handled by the conductor plugin, not the model itself.- Languages other than English and Portuguese are untested.
License
Apache 2.0 β same as the upstream laya checkpoint.
Model tree for mvilacad/laya-conductor
Base model
convaiinnovations/laya