Instructions to use Alogotron/a0-toolcall-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Alogotron/a0-toolcall-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct") model = PeftModel.from_pretrained(base_model, "Alogotron/a0-toolcall-lora") - Notebooks
- Google Colab
- Kaggle
a0-toolcall-lora (Qwen2.5-7B-Instruct)
QLoRA adapter teaching a 7B open model to reliably emit Agent Zero tool-call JSON — one object with thoughts, headline, tool_name, tool_args — instead of stalling in plain text. Motivated by two real stall incidents (2026-10-02/03) of a local 27B agent model.
Trained on Alogotron/a0-toolcall-sft (v2): 887 SFT records of real agent conversations (brain + arm instances), chat-level split so the 59 eval records come from 4 fully held-out chats.
Training
- QLoRA NF4 4-bit base + bf16 LoRA (r=16, alpha=32, dropout 0.05; all attention+MLP projections)
- 3 epochs, effective batch 8 (1×8 accum), lr 1e-4 cosine, warmup 5%, seq 1536, seed 42
- Prompt-masked loss (labels −100 on system+user context; loss only on target JSON)
- Single RTX 3090 24GB slot; runs alongside resident Ollama (no eviction; ~12GB/GPU free)
Eval: tool-call validity & stall rate (held-out chats)
Deterministic (greedy) generation, 512 new tokens, judged: valid = parses as JSON dict with all four keys, known tool name, dict tool_args; invalid = parses but violates shape; stall = no parseable JSON object.
| model | validity | invalid | stall | n |
|---|---|---|---|---|
| Qwen2.5-7B-Instruct base | 1.7% (1) | 98.3% (58) | 0% (0) | 59 |
| + this adapter | 84.7% (50) | 0% (0) | 15.3% (9)* | 59 |
All 9 tuned stalls are 512-token-cap truncations in a single held-out chat with unusually long tool_args (valid JSON starts, cut mid-object) - harness cap, not model failure. Base emitted 58 near-miss JSONs (invalid) and only 1 fully valid call. Full protocol + per-record details in eval report / eval-.json.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('Qwen/Qwen2.5-7B-Instruct', device_map='auto')
model = PeftModel.from_pretrained(base, 'Alogotron/a0-toolcall-lora')
# system prompt: instruct EXACTLY ONE JSON object with thoughts/headline/tool_name/tool_args
Intended use & limits
Use to harden small agent models against tool-call format stalls in Agent Zero-style loops. Trained on one agent household's conversations — tool distribution skews to code_execution (66%); JSON arguments generalize but niche tool semantics may not. Not a general-instruct tune.
- Downloads last month
- 14