turnercore's picture
Publish LFM2.5 Automaticity V9 response-only LoRA
0cbacae verified
|
Raw
History Blame Contribute Delete
2.91 kB
---
license: other
base_model: LiquidAI/LFM2.5-1.2B-Instruct
library_name: peft
pipeline_tag: text-generation
tags:
- function-calling
- tool-use
- automaticity
- automaticity-v9
- lora
- sft
- transformers
- trl
- unsloth
---
# LFM2.5 1.2B + Automaticity V9 LoRA
Rank-16 response-only LoRA trained for one epoch on the private Automaticity V9
friendly direct-tool corpus. This is the strongest current V9 validation
candidate, not a production-promoted autonomous router.
The model routes one current thought to at most one available tool, or makes no
tool call. Training used LFM2.5's native marked Python-call-list format and loss
only on the assistant turn.
## Training
- Base: `LiquidAI/LFM2.5-1.2B-Instruct`
- Base/tokenizer revision: `868df74dd56ff8a0c2ac5dbf281690c2dbebe4c9`
- Rows: 4,900; dataset SHA-256: `3fb79e5fe3cf762b3258c5674a806903e310aebb35d8ed153a525b0377b3bd8f`
- Context: 2,048 tokens; no truncation; maximum rendered row 1,978 tokens
- Precision: ROCm BF16 LoRA, not QLoRA
- LoRA: rank 16, alpha 16, dropout 0; `q/k/v/out/in_proj` and `w1/w2/w3`
- Epochs: 1; linear learning-rate schedule; 3% warmup
- Peak learning rate: 2e-4; weight decay: 0.001
- Effective batch: 16 (4 x 4 gradient accumulation)
- Seed: 3407
- Loss: native assistant response only
- Trainer runtime: 4,037 seconds
- Adapter SHA-256: `e81bdda7e1a684ae3a3f8d952303446d974c9ededf0cf383b8c76791112340ea`
## Frozen validation result
Evaluation used 1,050 private validation rows with normal five-tool retrieval,
no gold injection, 100% action-gold retrieval recall, and no decoding constraint.
The validation dataset SHA-256 is
`85094c96ca7fa2f96cbb0f7f85bd08510d56b9d4639646156d8806680bca9715`.
| Metric | Result |
| --- | ---: |
| End-to-end exact | 95.33% |
| Routing | 98.00% |
| Action exact | 84.94% |
| No-tool precision | 99.86% |
| No-tool recall | 99.73% |
| Argument schema validity | 99.90% |
| Listed-tool rate | 99.90% |
| Valid-call rate | 100% |
| Latency average | 1.242 s |
| Latency p50 | 0.518 s |
| Latency p95 | 4.520 s |
| No-tool latency average / p95 | 0.471 s / 0.661 s |
| Action latency average / p95 | 3.065 s / 8.413 s |
The untuned base on the identical ROCm validation condition scored 32.29%
end-to-end exact, 45.52% routing, 29.17% action exact, and 33.60% no-tool recall.
## Limitations
This adapter is not yet promoted for autonomous execution. The frozen validation
set still contains 20 wrong-tool rows, 28 wrong-argument rows, and one unlisted
call. Action p95 latency is 8.413 seconds, 6.3% slower than the untuned action
p95 even though aggregate latency improved substantially. Use strict listed-name
and schema validation or constrained decoding and reject invalid calls at runtime.
Constraints cannot repair semantically wrong listed tools or schema-valid wrong
arguments.
The private dataset and row-level evaluation repository is
`turnercore/automaticity-v9`.