chendren/cx-decider-applied-v1

Verdict: effective on applied CX decisions, not noise. 0.9417 on 240 held-out rows (12 per task, unseen in training). Complements β€” does not replace β€” the loop gate below.

20 applied CX decisions for 2026 practices: intent, channel, category, human-handoff, self-serve, escalation, churn, complaint, approval, fraud, sentiment, priority, csat-risk, VIP, callback, PII, translation, repeat-contact, SLA-risk, next-best-action. Each with calibrated confidence. Continual-tuned from chendren/cx-decider-2b-v1's starting point on a Mac.

Why two models, not one (read this first)

A joint checkpoint covering all 24 decisions was trained and rejected: adding the 20 applied tasks dragged loop safe_to_log 1.0000 β†’ 0.7819 (the veto signal β€” the one thing that must not break) while applied tasks scored 0.955. Anchor-retraining with 3Γ— loop upweighting moved loop only to 0.8831 while distorting choice calibration (ECE 0.108). Multi-task interference in the shared head, not a grind problem β€” more steps pick a point on the tradeoff, not a way off it.

So the architecture is two specialists from one lineage:

  1. chendren/cx-decider-2b-v1 β€” the loop gate (4 decisions: next_tool, safe_to_log, needs_case, urgency). Published, proven, untouched. Use it for gating.
  2. This repo β€” the applied head (20 decisions below). Same v1 starting point, 135 steps on applied rows only. Use it for triage, routing, risk flags.

Both load with strands_decider.infer.load_engine. The failure pattern that forced the split: applied tasks (all state-text judgments like fraud/sentiment) rewarded reading words; loop safe_to_log requires reading state ("create case" mentioned in the request but absent from completed tools). The shared head shifted toward words. Documented so nobody retries the joint run blind.

Try it

from strands_decider.infer import load_engine
eng = load_engine("chendren/cx-decider-applied-v1", device="mps")  # cuda | mps | cpu
r = eng.ask(
    "Maya R. via live chat requests her 6th refund in 3 weeks for $400 headphones, each claimed as never arrived though carrier shows delivered.",
    {"fraud": {"type": "noul", "instructions": "This shows signs of fraud or account takeover."},
     "intent": {"type": "choice", "instructions": "What is the customer intent?",
                "criteria": {"billing": "money, invoices, charges", "technical": "bugs, outages",
                             "account": "login, profile", "order_status": "where is my order",
                             "product_info": "product questions", "cancellation": "cancel service",
                             "complaint": "unhappy, formal complaint", "other": "anything else"}}})
# fraud -> ~0.95 true; intent -> complaint/billing at high confidence

Full rubric (instructions, options, levels): chendren/cx-decisions-v2-data. Use the exact question text there β€” paraphrased probes underperform the rubric (measured during testing), so the numbers below describe rubric-worded questions.

Training (exact)

  • Init: checkpoints/cx-decider-v1-mac (the published loop gate's starting lineage) β€” kept LoRA + head, adapters + head trainable (17,872,384 / 1,899,697,472, 0.94%). Temperatures reset, refit after.
  • Data: cx_applied_train.jsonl, 2,160 rows (90% of 2,400; 108 per each of 20 tasks). LLM-inferred states + judged labels, even score spreads, 2 instruction variants each. Val: cx_applied_val.jsonl, 240 rows (12/task), unseen in training.
  • Run: 135 optimizer steps, micro-batch 8 Γ— grad-accum 2, head LR 5e-4, LoRA LR 2.5e-5, AdamW Ξ²=(0.9, 0.95), cosine, warmup 3%, clip 1.0, head-only decay 0.01, max_length 1024, seed 0, KL 0.3 every 2nd step. Loss 0.403 β†’ 0.255 (trace: history.json).
  • Machine: Apple M2 Max, MPS, torch 2.14.1.
  • Calibration: refit on 400 held-out applied rows β€” global 0.4675, noul 0.2784. ECE 0.0838 β†’ 0.0413, acc 0.9675 on fit set.

Test results (measured, MPS)

eval result
Applied val, overall (n=240, 12/task unseen) acc 0.9417, ECE 0.0703
Per-task highs (1.000) callback, category, complaint, nba, repeat, sla, translate
Per-task lows priority 0.767, channel 0.842, churn/vip ~0.86–0.91, escalation 0.900
Loop-heldout via THIS checkpoint (n=4,328) acc 0.8953 β€” do not use it for gating; v1 holds 0.9499

Weak spots are real: cx_priority (0.767) and cx_channel (0.842) confuse adjacent levels/options. Same 0.9 act-bar as v1 applies β€” mid-band answers need confirmation.

When NOT to use it

  • Loop gating β€” use v1 (cx-decider-2b-v1). This head's loop accuracy (0.8953) trails v1 (0.9499) and its safe_to_log is untested for veto duty.
  • Paraphrased questions β€” measured underperformance vs rubric wording; stay on the dataset's instruction text.
  • Priority/channel below 0.9 β€” the two weakest tasks; confirm or route to human.
  • Non-CX decisions β€” generic retention unmeasured (same caveat as v1).
  • Commercial use blind β€” applied states are invented (no PII), but v1's bitext lineage note applies to the shared starting point.

Files

  • head.safetensors + lora/ β€” weights (use with Qwen/Qwen3.5-2B-Base)
  • strands_decider_config.json β€” architecture + fitted temperatures
  • train_config.json / history.json / provenance.json β€” run record
  • tokenizer.json, tokenizer_config.json, chat_template.jinja β€” serving files
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for chendren/cx-decider-applied-v1

Adapter
(33)
this model