aliangdw's picture
FFT finetune of Robometer-4B on both so101 datasets (armnet + molmoact)
89f8382 verified
|
Raw
History Blame Contribute Delete
1.05 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3-VL-4B-Instruct
tags:
  - robometer
  - reward-model
  - rbm
  - armnet
  - molmoact
library_name: transformers

Robometer-4B FFT finetuned on both so101 datasets (Armnet so101 + MolmoACT2 so101)

Full fine-tune (FFT, no LoRA) of Robometer-4B on both SO-101 datasets:

  • Armnet benchmark so101 (villekuosmanen_armnetbench_robometer_v01_so101)
  • MolmoACT2 so101 (ykorkmaz_molmoact2_so100_101_rbm_molmoact2_so100_101)

Qwen3-VL-4B backbone, 1500 steps on 4x H200. This is the best checkpoint (step 750).

Key result: training on both so101 datasets improves generalization

Metric armnet-only finetune this (both so101)
Armnet so101 reward-alignment Pearson 0.766 0.782
Armnet so101 policy-ranking Kendall 0.973 0.94
Molmoact so101 reward-alignment Pearson (held-out) 0.751 0.902

Training on both datasets raised molmoact so101 Pearson from 0.75 → 0.90 while keeping armnet performance roughly intact.