YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Act-PRM SFT LoRA checkpoints โ tau2 retail
LoRA adapters (r8_a16, base Qwen/Qwen3-4B-Instruct-2507) from Act-PRM supervised fine-tuning on tau2-bench retail. Variants: actions_only (baseline), expert_thoughts (oracle), thoughts_{policy,base}[_last] (Act-PRM inferred thoughts), each in hide-observations and full-context regimes. Adapters live under retail/<variant_regime>/.
Held-out action-only eval (lower PPL / higher acc = better next-action fit)
| variant_regime | action-only PPL | action-acc |
|---|---|---|
| thoughts_base_last_heldout_fullctx | 2.982 | 0.803 |
See the project for methodology (Act-PRM: infer latent thoughts behind action-only demos via offline EM).
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support