Nima

A small decision head that teaches frozen Qwen to choose among textual tool schemas.

Nima takes a request and a supplied set of tool descriptions, then returns one score per candidate and a selected identity. It bypasses autoregressive generation: no JSON arguments or reasoning tokens are generated, and no tools are executed.

This release contains the Joint-Schema V2 decision head, with 1,318,657 parameters, plus standalone MLX inference code. The Qwen3-1.7B 4-bit backbone remains completely frozen and is downloaded separately at its pinned revision. This is not a merged Qwen checkpoint or a LoRA adapter.

Author: Anand Mohanan / neimand.

How it works

Context + request + all candidate schemas
                    ↓
       Frozen Qwen3-1.7B (4-bit)
                    ↓
     Final normalized hidden states (2048D)
                    ↓
   Shared LayerNorm + projection (256D)
                    ↓
 Candidate β†’ request/context cross-attention
                    ↓
   Bidirectional joint-schema Transformer
                    ↓
    Shared scalar score per candidate
                    ↓
           Softmax β†’ decision

Each candidate representation is the mean of Qwen states over its complete schema span. Evidence consists of individual context/request token states. The head uses four attention heads, one pre-normalized Transformer layer, a 512-dimensional FFN and no dropout or candidate-index embeddings. Padding masks support variable candidate counts; there is no fixed Linear(8) classifier.

The head is permutation-equivariant on fixed feature vectors. The complete pipeline is order-sensitive: reordering textual schemas changes causal Qwen states. Dynamic inputs do not establish unseen-domain reasoning ability.

Run locally

Use an Apple Silicon environment supporting MLX; the recorded experiment used Python 3.12, MLX 0.32.3 and MLX-LM 0.32.0. This custom architecture is not loaded through AutoModel or a standard chat pipeline, and this release does not provide a hosted inference service.

Download the release:

python -m pip install 'huggingface-hub>=1.0,<2.0'
hf download neimand/nima --local-dir nima
cd nima
python -m pip install -r requirements.txt
python -B inference.py --input example.json

The final command downloads and verifies the pinned Qwen backbone before predicting. Add --local-files-only after the backbone has been cached to require offline loading. Weight verification and model loading are excluded from the returned request latency.

For your own request, provide JSON with:

  • prompt: the request text.
  • context: an object containing current_date and timezone.
  • choices: at least two uniquely named schemas, each with name, description, and a parameters object. Include an explicit no_tool option.

example.json is an input illustration, not a newly measured result. Output contains prediction, candidate-keyed logits and probabilities, token count and latency. Probabilities are uncalibrated and depend on the supplied candidate set. Inputs over 1,536 tokens are rejected without truncation. Serialization and pooling are preserved from V2; do not silently replace them with a chat template.

Training

Only the head was trained, using cached frozen-Qwen features. The dataset covered eight familiar capabilities: arithmetic, weather, web search, private-document retrieval, structured records, calendar lookup, email sending and abstention. Wording, schema aliases, candidate counts and order varied; scenario/template families were separated across splits.

Setting Value
Training / selection / calibration / confirmation requests 512 / 128 / 128 / 256
Effective contextual training views 1,706
Optimizer AdamW
Learning rate / weight decay 0.0003 / 0.01
Batch size / epoch budget / seed 8 / 80 / 42
Checkpoint selection Minimum selection cross-entropy; selected epoch 1
Temperature 1; no fitting or confidence threshold
Trainable backbone parameters 0

The training script generated up to four distinct candidate permutations per training request; sets with fewer unique permutations contribute fewer views. Selection used two orders. Calibration data was separate but no temperature fitting was performed.

training_protocol.json, training_metadata.json and dataset_manifest.json preserve settings, recorded checks, split counts and hashes. Raw datasets and full training infrastructure are not bundled in this inference release.

Evaluation

These are adapted bounded-choice routing diagnostics, not official leaderboard scores. No public-benchmark training, model selection, thresholds or temperature fitting was performed.

Panel, canonical order Correct Accuracy
ToolRet APIGen tool selection 75/96 78.1%
BFCL function-name selection 89/100 89.0%
BFCL irrelevance: correctly selected no_tool 42/100 42.0%
Original V2 family-separated confirmation 200/256 78.1%
Later V3 paired confirmation, unchanged V2 checkpoint 312/416 75.0%

ToolRet supplied one judged-positive tool, two lexical-hard distractors, one seeded-random distractor and no_tool. This is not full-corpus retrieval, and unjudged distractors can also be relevant. BFCL used single-message, single-call selection and irrelevance cases; arguments, execution, parallel calls and multi-turn behavior were not scored. Full schemas were retained; a predeclared length-compatible subset excluded over-limit cases across all tested orders.

Known failure modes

  • Abstention: 58/100 BFCL irrelevance cases incorrectly selected a function. Across the two BFCL panels, no_tool precision was 100%, recall 42%, with zero canonical false abstentions among supported requests. No actual tools execute.
  • Candidate order: predictions changed across tested orders on 37/96 ToolRet requests (38.5%), 23/100 BFCL selection requests, and 29/100 BFCL irrelevance requests. Up to four orders were used; two-choice sets have two. ToolRet accuracy ranged from 65.6% to 78.1% across orders.
  • Generalization: templated single-author training data and correlated examples limit conclusions. Public-data pretraining overlap is unknown. Familiar-capability confirmation sets do not demonstrate unseen-domain performance.
  • Output scope: routing only, without argument generation, execution safeguards or a complete agent workflow. Treat it as a research checkpoint, not a reliable autonomous executor.

Recorded median canonical latency was 0.680s for ToolRet, 0.566s for BFCL selection and 0.373s for BFCL irrelevance, including tokenization, one frozen-backbone forward, the head and synchronization. Loading, serialization and extra order diagnostics were excluded. Historical timings are not controlled speed comparisons. Peak MLX allocation was approximately 1.66 GB, which is not total system memory.

See evaluation_results.json for the saved metrics and artifact hashes. The release was checked for tensor shapes/counts, weight integrity, Python syntax and encoding parity without rerunning training or inference. The standalone loader has not received a new end-to-end inference run for this release.

Downloads last month
-
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for neimand/nima

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(2)
this model