Instructions to use srpone/instinct-tuned-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use srpone/instinct-tuned-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="srpone/instinct-tuned-4b")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("srpone/instinct-tuned-4b") model = AutoModelForMultimodalLM.from_pretrained("srpone/instinct-tuned-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ZooWork Instinct Tuned 4B
ZooWork · Instinct homepage and hosted API · API model ID: instinct-tuned-4b
Instinct Tuned 4B is a 4B-parameter decision model. You give it a shared state (a message, a document, or a JSON object) and a question with a fixed set of candidate answers. It returns a calibrated probability for each candidate from one forward pass. It does not generate text.
Do not use
generate()or atransformerspipeline with this model. The decision is read from the label-token logits with a fixed prompt and temperature; use the reference runtime below.
It supports three question types:
| Type | Question | Output |
|---|---|---|
choice |
Pick one of 2–16 described options | a probability per option, the argmax, and a confidence |
noul |
Is this proposition true of the state? | P(yes) |
score |
Place the state on an ordered scale of 2–16 levels | a probability per level, and the expected level |
The model is a LoRA fine-tune, merged into the base weights, of Qwen/Qwen3.5-4B.
Quick start
The readout is not a standard generate() or classification head, so use the reference runtime: SerendipityOneInc/instinct.
git clone https://github.com/SerendipityOneInc/instinct && cd instinct
pip install -e ".[gpu]"
instinct-decide --model instinct-tuned-4b models/instinct-tuned-4b/examples/request.json
from instinct import InstinctModel, answer
model = InstinctModel.from_pretrained("instinct-tuned-4b") # downloads srpone/instinct-tuned-4b
print(answer(model, {
"state": "Where is my package? I ordered it last week and it still hasn't arrived.",
"questions": {
"intent": {"type": "choice", "instructions": "Which intent does the message express?",
"criteria": {"track_order": "Wants to know where an order is",
"cancel_order": "Wants to cancel an order",
"billing_question": "Asks about a charge or payment"}},
"angry": {"type": "noul", "instructions": "The customer is angry."},
},
}))
How it works
- The request is rendered with the fixed prompt
instinct.prompt.v1:Shared state:\n<state>\n\n- then a sorted-key JSON object
{"criteria": [{"description", "label"}], "instructions", "primitive"}, with the candidates labelledA,B, … - then
Return only the selected letter: A, B, ….\nAnswer:
- The prompt goes through the chat template with
add_generation_prompt=Trueandenable_thinking=False. - The last hidden state at full depth (layer 32) passes through the final norm. Only the LM-head rows of the candidate label tokens are applied ("candidate-row readout").
- The logits are divided by the serving temperature T = 2.80, then softmaxed. T is stored in
decision_config.json.
A noul question is scored as a two-way choice between the canonical candidates yes ("The stated proposition is true.") and no. Custom true/false wording is accepted but not shown to the model. A score question is scored as a choice over levels "0"…"n-1". A structured state is rendered with json.dumps(state, ensure_ascii=False).
Evaluation
The table reports JevBench public (231 items; 198/231 overall), full depth, bf16, scored with the official JevBench client, items in their original option order:
| Scope | Correct | Accuracy |
|---|---|---|
| All published tasks | 198/231 | 85.71% |
| easy | 48/48 | 100.00% |
| standard | 69/72 | 95.83% |
| hard | 81/111 | 72.97% |
These public items were also used during our development for model selection, so treat the table as a reference point, not a held-out score.
Calibration: the expected calibration error on the hard split after temperature scaling is 0.04–0.11, depending on option order (see Limitations).
| Serving latency | Value |
|---|---|
| Direct-upstream p50 | 62 ms |
| Direct-upstream p95 | ≈110 ms (estimated) |
Latency refers to warmed, serial serving and excludes public Internet, TLS and gateway overhead. The p50 is measured; the p95 is an estimate from the measured direct-upstream p50 and observed end-to-end tail spread, not an SLO.
Training
- Base: Qwen3.5-4B, text path only. The vision tower weights are carried over from the base unchanged.
- Method: LoRA (r = 32, α = 64) on the decision objective (cross-entropy over the candidate label tokens), with auxiliary readouts at layers 16 and 24. The auxiliary readouts are not used at inference. Settings: 1 epoch, learning rate 1e-4, effective batch 32. The LoRA was merged in fp32 and saved in bf16.
- Data: about 9.4k decision items: a base decision training set (~5.9k), plus ~900 synthetic hard examples blind-verified by an independent model and upsampled 4x. The training data is not released at this time; we plan to release it.
- Decontamination (training data only): every training item has less than 20% 13-gram overlap with the JevBench public set and with our held-out development sets. This check covers the training data; the public set itself was used for model selection (see Evaluation).
- Temperature: T was fitted on a held-out development set drawn from the same task distribution as the benchmark.
Limitations
- Option-order sensitivity. Probabilities depend on which candidate gets which letter. Accuracy is stable, but calibration error on hard items moves by 0.04–0.07 between orders. The runtime uses the order given in the request.
- Text only, 8,192-token limit. Longer inputs are rejected, not truncated. Image and video inputs are not supported.
- Mostly English evaluation.
- T was fitted on a held-out development set from the same task distribution as the benchmark, so the probabilities are calibrated for that distribution. Re-fit T on your own labelled data if your domain differs.
- The model makes decisions; it does not explain them.
Files
| File | Purpose |
|---|---|
model-0000{1,2,3}-of-00003.safetensors |
bf16 weights |
model.safetensors.index.json |
shard index |
config.json, generation_config.json |
Qwen3.5 config. The top-level dtype is bfloat16; text_config and vision_config say float32 because that was the merge precision. The shards are bf16. |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
tokenizer and chat template |
decision_config.json |
prompt version, serving temperature, readout depth |
SHA256SUMS |
file hashes |
License
The weights are released under Apache-2.0, the license of the base model Qwen3.5-4B.
About
Instinct is developed by ZooWork. Learn more at instinct.zoowork.ai.
- Downloads last month
- -