LBM v3.6 DPO

This is a merged, vLLM-ready LBM v3.6 checkpoint specialized for the Thai market. It starts from amityco/lbm-v3.6 and applies same-decision DPO on fresh, externally judged Thai customer rollouts.

Output contract

First-person Rationale -> Acceptance Probability -> Final Decision: ACCEPT/REJECT

Frozen 2,000-row Thai diagnostic

Task ROC AUC
Coupon response 0.6811
Promotion response 0.6583
New-product adoption 0.7271

Gate pass: True.

Scope and limitations

  • Behavior events use customer-held-out Lotus-derived ground truth and opportunity proxies.
  • The small diagnostic is outcome-balanced and is not a population calibration estimate.
  • HBA/WTP remain evaluation-only. v3.6 post-training does not yet dominate older LBM checkpoints on every survey metric, so run the complete frozen Thai benchmark before deployment.
  • Do not interpret predicted acceptance as causal coupon lift.
Downloads last month
14
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for amityco/lbm-v3.6-dpo

Base model

amityco/lbm-v3.6
Finetuned
(1)
this model