AutoScientist Competition โ€” Other Model

Qwen2.5-0.5B-Instruct adapted for other via Adaption Labs AutoScientist v5: 4-bit QLoRA SFT (r=32, alpha=64) then DPO (beta=0.1) on chosen/rejected pairs.

  • DPO reward accuracy: 0.8181818181818182
  • DPO reward margin: 4.244321004910902

Dataset: Rishidar/autoscientist-other-dataset. Also mirrored on Kaggle: rishidard/autoscientist-other-qlora.

Downloads last month
4
Safetensors
Model size
0.5B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Rishidar/autoscientist-other-qlora

Finetuned
(949)
this model

Dataset used to train Rishidar/autoscientist-other-qlora