StockForge question router (Qwen3-0.6B, LiteRT INT4)

The on-device model behind the StockForge inventory app's assistant. It turns a shop owner's question in English, Spanish or Portuguese into one lookup, as one line of JSON:

¿qué vendió Carlos ayer?  ->  {"intent": "stock_movements", "args": {"hours": 48, "person": "Carlos", "reason": "sale"}}

It never answers with data itself: the app sends the lookup to its server, which answers from the user's own records.

  • Base: Qwen/Qwen3-0.6B, full fine-tune on ~1.8k synthetic examples
  • Format: .litertlm for LiteRT-LM, dynamic INT4 (block 32), 1280-token context
  • System prompt: Route the shop inventory question to one lookup. Reply with one JSON line. (thinking off)
  • Held-out test (120 questions, 40 per language): 98% correct lookup, 95% lookup + arguments; median 1.1 s on a phone CPU

License: Apache 2.0, as the base model.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for queensone/stockforge-router

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1342)
this model