mai-nano-router-2b

A Qwen3.5-2B-Base LoRA SFT router, fine-tuned for intent classification + slot extraction as part of the myMAI Spotlight project's finetuning/ pipeline (en/ko/ja/zh).

This is the larger sibling of LUNAV/mai-nano-router (Qwen3.5-0.8B-Base) โ€” recommended for machines with 16GB+ RAM, where the extra headroom is worth the accuracy/robustness gain.

Results (test set, 563 examples)

Metric Value
JSON validity 100%
Intent accuracy 98.93%
Slot F1 0.789
Per-language intent accuracy en 97.9% / ja 99.3% / ko 98.6% / zh 100%

Verified against the real production inference path (llama-server, Q8_0 GGUF) โ€” not just the training-time HF/PyTorch backend.

Files

  • mai-nano-router-2b-q8_0.gguf (Q8_0, ~1.9GB) โ€” the quantization myMAI Spotlight ships.
  • mai-nano-router-2b-f16.gguf (F16, ~3.6GB) โ€” unquantized reference.

Usage

Same runtime contract as mai-nano-router: served via llama.cpp's llama-server, prompted with the canonical system prompt from this project's dataset/final/chat/*.jsonl, expecting a single-line JSON completion ({"intent": ..., "confidence": ..., "slots": {...}}).

Trained and converted from cornch-k/mymai's finetuning/ pipeline.

Downloads last month
-
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LUNAV/mai-nano-router-2b

Quantized
(25)
this model