a0-toolcall-lora (Qwen2.5-7B-Instruct)

QLoRA adapter teaching a 7B open model to reliably emit Agent Zero tool-call JSON — one object with thoughts, headline, tool_name, tool_args — instead of stalling in plain text. Motivated by two real stall incidents (2026-10-02/03) of a local 27B agent model.

Trained on Alogotron/a0-toolcall-sft (v2): 887 SFT records of real agent conversations (brain + arm instances), chat-level split so the 59 eval records come from 4 fully held-out chats.

Training

  • QLoRA NF4 4-bit base + bf16 LoRA (r=16, alpha=32, dropout 0.05; all attention+MLP projections)
  • 3 epochs, effective batch 8 (1×8 accum), lr 1e-4 cosine, warmup 5%, seq 1536, seed 42
  • Prompt-masked loss (labels −100 on system+user context; loss only on target JSON)
  • Single RTX 3090 24GB slot; runs alongside resident Ollama (no eviction; ~12GB/GPU free)

Eval: tool-call validity & stall rate (held-out chats)

Deterministic (greedy) generation, 512 new tokens, judged: valid = parses as JSON dict with all four keys, known tool name, dict tool_args; invalid = parses but violates shape; stall = no parseable JSON object.

model validity invalid stall n
Qwen2.5-7B-Instruct base 1.7% (1) 98.3% (58) 0% (0) 59
+ this adapter 84.7% (50) 0% (0) 15.3% (9)* 59

All 9 tuned stalls are 512-token-cap truncations in a single held-out chat with unusually long tool_args (valid JSON starts, cut mid-object) - harness cap, not model failure. Base emitted 58 near-miss JSONs (invalid) and only 1 fully valid call. Full protocol + per-record details in eval report / eval-.json.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('Qwen/Qwen2.5-7B-Instruct', device_map='auto')
model = PeftModel.from_pretrained(base, 'Alogotron/a0-toolcall-lora')
# system prompt: instruct EXACTLY ONE JSON object with thoughts/headline/tool_name/tool_args

Intended use & limits

Use to harden small agent models against tool-call format stalls in Agent Zero-style loops. Trained on one agent household's conversations — tool distribution skews to code_execution (66%); JSON arguments generalize but niche tool semantics may not. Not a general-instruct tune.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alogotron/a0-toolcall-lora

Base model

Qwen/Qwen2.5-7B
Adapter
(2800)
this model

Dataset used to train Alogotron/a0-toolcall-lora