jev-decisions-v2-model

Jev/Tev1-style single-letter decision model: a full fine-tune of Qwen/Qwen3.5-0.8B on GhostScientist/jev-decisions-v1 (8,702 train rows). Given a structured state / question / options task, it answers with exactly one option letter.

GGUF exports for Ollama / llama.cpp: GhostScientist/jev-decisions-v2-model-gguf (Q4_K_M + Q8_0, with a ready-to-use Modelfile).

Measured results (dev split, n=400, greedy, max_new_tokens=8)

eval v1 (jev-decisions-v1-model) v2 (this model)
accuracy, a10g GPU bf16 0.7525 0.8000
valid-letter rate 1.000 1.000
latency mean/p50/p95 (a10g GPU, ms) 139 / 132 / 164 138 / 129 / 158
accuracy, llama.cpp Q4_K_M, 4 threads, n=200 0.740 0.820
valid-letter rate, CPU 1.000 1.000
latency mean/p50/p95 (CPU, ms) 1404 / 1320 / 2528 1403 / 1332 / 2467

By-source (GPU eval): strongest seed_banking 0.979, seed_agnews 0.839, seed_boolq 0.804; weakest seed_sst5 0.432, synthetic_scheduling_decision 0.500 (n=8), seed_nli 0.677.

Serving notes (important)

The Qwen3.5 chat template defaults to thinking mode. This model was trained on bare-letter completions, so where the letter lands depends on the server:

  • llama-server: can route the letter to reasoning_content (empty content). Prefer chat_template_kwargs: {"enable_thinking": false} — measured there, accuracy went 0.00 → 0.82 on the same weights.
  • Ollama: the letter lands in message.thinking. Pass "think": false in the /api/chat payload, or fall back to thinking when content is empty. Verified: valid rate 1.00, accuracy 0.925 on 40 samples.
  • A robust client reads content, then thinking/reasoning_content, then applies the letter regex to whichever is non-empty.

GGUF conversion requires convert_hf_to_gguf.py --no-mtp (Qwen3.5 declares an MTP layer it has no tensors for; default conversion produces a 25-block metadata header over 24 blocks and llama-server refuses to load it).

Training

Full-parameter SFT with TRL 1.14.1 / transformers 5.18.0, conversational prompt-completion format, loss on the completion only (letter + EOS): lr 2e-5, cosine, warmup 0.03, effective batch 8 (4 x 2), 2 epochs, max_len 2048, bf16 + gradient checkpointing, on a10g-small (~77 min). Trackio: jev-decisions-v2-sft-trackio.

Final: eval_loss 0.1536, token accuracy 0.9423.

Citations

Cite TRL as:

@software{vonwerra2020trl,
  title   = {{TRL: Transformers Reinforcement Learning}},
  author  = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
  license = {Apache-2.0},
  url     = {https://github.com/huggingface/trl},
  year    = {2020}
}
Downloads last month
113
Safetensors
Model size
0.9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GhostScientist/jev-decisions-v2-model

Finetuned
(479)
this model

Dataset used to train GhostScientist/jev-decisions-v2-model