diffcider-typed-decisions

A typed-decision model made with sysone: the encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 (lora) under a 0-layer yesno head, trained on LocalLLaMA/typed-decisions (all) at d0e2f0c4.

The adapter makes the masked diffusion language model Qwen3-0.6B-diffusion-mdlm-v0.1 a typed-decision model: it answers choices, scores and yes/no statements in one forward pass, reading the model's own Yes against its No at a mask after each option, in jev-dllm's layout. On typed-decisions' 2,000 test decisions it scores 0.768, where the model scores 0.409 untrained and jev-dllm reports 0.5225 for a full fine-tune of it. With the adapter switched off the model is the published one exactly, and it writes (dec.generate(prompt)). sysone's mdlm module loads the checkpoint's a2d-qwen3 model type with its own classes, so no remote code runs.

The repo holds the head and a LoRA adapter, not the encoder's weights: Decider.load downloads dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 at commit c8d24a3f4a from the Hub and puts the adapter on it.

Use it

Install sysone from GitHub, in Python 3.11 or newer:

pip install git+https://github.com/sgaseretto/sysonelib

A state is whatever the decision is about, as JSON-like data. Each question has a type: a choice among named options, a score on an ordered scale, or a noul, a statement that is true or false. The options are named when asking, so they can be new ones.

from sysone.inference import Decider

dec = Decider.load("sgaseretto/diffcider-typed-decisions")

state = {
    "account": {"tier": "premium", "tenure_months": 26, "prior_tickets_90d": 2},
    "thread": [
        {
            "role": "customer",
            "text": "I was charged twice for my March invoice, and nobody has answered my last two emails.",
        },
        {"role": "agent", "text": "I'm sorry about that. Could you share the invoice number so I can look into it?"},
        {"role": "customer", "text": "It's INV-3381. Refund the second charge today or I'm cancelling the account."},
    ],
}
questions = {
    "category": {
        "type": "choice",
        "instructions": "What is this customer conversation primarily about?",
        "criteria": {
            "billing": "A charge, invoice, subscription or payment problem.",
            "refund": "The customer is explicitly asking for money back.",
            "technical": "The product or service is not working as expected.",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How time-sensitive is this conversation?",
        "criteria": [
            "No time pressure; can wait indefinitely.",
            "Routine; handle within the normal queue.",
            "Elevated; should be handled within the same week.",
            "Critical; requires action within the same day.",
        ],
    },
    "needs_human": {
        "type": "noul",
        "instructions": "This conversation requires a human agent rather than automated handling.",
    },
}
answers = dec.predict(state, questions)

answers holds an answer per question, in the Jev schema: a choice names the likeliest option and gives every option's probability, a score gives its expected level on the scale (0 for the first) and every level's probability, and a noul gives the probability that its statement is true; each says how confident it is. This model's answers to the example:

{
    "category": {
        "type": "choice",
        "choice": "refund",
        "probabilities": {"billing": 0.429, "refund": 0.5677, "technical": 0.0033},
        "confidence": 0.3597,
        "answer_confidence": 0.5677,
    },
    "urgency": {
        "type": "score",
        "score": 2.6442,
        "legend": {
            "0": "No time pressure; can wait indefinitely.",
            "1": "Routine; handle within the normal queue.",
            "2": "Elevated; should be handled within the same week.",
            "3": "Critical; requires action within the same day.",
        },
        "probabilities": {"0": 0.0069, "1": 0.0483, "2": 0.2387, "3": 0.7062},
        "confidence": 0.4459,
        "answer_confidence": 0.7062,
    },
    "needs_human": {"type": "noul", "noul": 0.7083, "confidence": 0.7083, "answer_confidence": 0.7083},
}

Results

On LocalLLaMA/typed-decisions's test split (2,000 decisions), calibrated, on a Kaggle T4:

metric value
accuracy 0.7675
accuracy_choice 0.7517
accuracy_score 0.7238
accuracy_noul 0.8417
ece 0.1496
brier 0.0536

On the validation split (400 decisions), during training:

metric value
loss 0.9801
accuracy 0.6925
soft_accuracy 0.4572
brier 0.0621
kl 0.1006
tv 0.167
ece 0.1495
score_mae 0.2234
within_one 1
accuracy_choice 0.7
accuracy_score 0.6062
accuracy_noul 0.85

Calibration temperatures: choice 0.945, choice:3-5 0.945, score 1.02, score:3-5 1.02, noul 1.02, noul:2 1.02.

Training

Data LocalLLaMA/typed-decisions (all) at d0e2f0c4; 1080 train, 80 valid, 120 calib cases
Encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1, lora (LoRA r=16, α=32, on down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj)
Head 0 layers, yesno scorer, readout anchor
Loss soft_ce_rps=0.25, preset t4, seed 0
Fit 1 1 epoch, cosine, lr 0.001 (encoder 0.0002), batch 8×4, fp16 on cuda: 300 steps in 87.4 min, peak memory 7.69 GB, final loss 0.878
Machine Intel(R) Xeon(R) CPU @ 2.00GHz, 4 cores, 31.3 GB; accelerator cuda (Tesla T4, Tesla T4); a kaggle run (job mdlm-lora-check)

Reproduce

  • sysone: sysonelib at commit 123ad8bac6 (main), with uncommitted changes in 22 files (diff SHA-256 77eb5a356ea0)
  • Ran: run_b.py
  • Python 3.13.15, sysone 0.3.0, torch 2.11.0+cu128, transformers 5.16.1, peft 0.20.0, accelerate 1.14.0, datasets 4.8.5, huggingface_hub 1.29.0, tokenizers 0.23.1, safetensors 0.8.0, numpy 2.1.3, fastcore 2.2.32, plum-dispatch 2.10.1; environment.txt lists every package
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sgaseretto/diffcider-typed-decisions

Finetuned
Qwen/Qwen3-0.6B
Adapter
(5)
this model

Dataset used to train sgaseretto/diffcider-typed-decisions

Evaluation results