UTI_L3_1000steps_1e5rate_03beta_CSFTDPO

This model is a fine-tuned version of tsavage68/UTI_L3_1000steps_1e5rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0069
  • Rewards/chosen: 2.2757
  • Rewards/rejected: -15.6836
  • Rewards/accuracies: 0.9900
  • Rewards/margins: 17.9593
  • Logps/rejected: -115.4733
  • Logps/chosen: -24.8934
  • Logits/rejected: -1.4719
  • Logits/chosen: -1.4307

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 2
  • eval_batch_size: 1
  • seed: 42
  • gradient_accumulation_steps: 2
  • total_train_batch_size: 4
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss Rewards/chosen Rewards/rejected Rewards/accuracies Rewards/margins Logps/rejected Logps/chosen Logits/rejected Logits/chosen
0.0 0.6667 50 0.0072 1.8402 -13.3590 0.9900 15.1992 -107.7247 -26.3451 -1.4305 -1.3941
0.0173 1.3333 100 0.0071 1.8455 -14.4051 0.9900 16.2506 -111.2116 -26.3273 -1.4331 -1.3960
0.0347 2.0 150 0.0069 2.3483 -14.9050 0.9900 17.2533 -112.8780 -24.6513 -1.4557 -1.4154
0.0 2.6667 200 0.0069 2.3179 -15.0160 0.9900 17.3339 -113.2480 -24.7526 -1.4584 -1.4180
0.0173 3.3333 250 0.0069 2.3120 -15.0851 0.9900 17.3971 -113.4783 -24.7723 -1.4616 -1.4212
0.0347 4.0 300 0.0069 2.3109 -15.1144 0.9900 17.4254 -113.5761 -24.7759 -1.4624 -1.4219
0.0173 4.6667 350 0.0069 2.3085 -15.1859 0.9900 17.4944 -113.8144 -24.7841 -1.4649 -1.4242
0.0173 5.3333 400 0.0069 2.2984 -15.2571 0.9900 17.5555 -114.0517 -24.8176 -1.4668 -1.4260
0.0173 6.0 450 0.0069 2.2945 -15.3467 0.9900 17.6412 -114.3504 -24.8307 -1.4680 -1.4272
0.0347 6.6667 500 0.0069 2.2859 -15.4295 0.9900 17.7154 -114.6264 -24.8593 -1.4694 -1.4284
0.0 7.3333 550 0.0069 2.2833 -15.5057 0.9900 17.7890 -114.8804 -24.8681 -1.4703 -1.4293
0.0347 8.0 600 0.0069 2.2775 -15.5762 0.9900 17.8538 -115.1155 -24.8872 -1.4709 -1.4298
0.0 8.6667 650 0.0069 2.2759 -15.6206 0.9900 17.8965 -115.2633 -24.8928 -1.4712 -1.4301
0.0173 9.3333 700 0.0069 2.2757 -15.6425 0.9900 17.9182 -115.3363 -24.8933 -1.4714 -1.4302
0.0 10.0 750 0.0069 2.2743 -15.6650 0.9900 17.9392 -115.4112 -24.8982 -1.4717 -1.4305
0.0173 10.6667 800 0.0069 2.2739 -15.6785 0.9900 17.9524 -115.4563 -24.8992 -1.4719 -1.4307
0.0 11.3333 850 0.0069 2.2703 -15.6667 0.9900 17.9370 -115.4169 -24.9113 -1.4717 -1.4306
0.0 12.0 900 0.0069 2.2749 -15.6771 0.9900 17.9520 -115.4516 -24.8959 -1.4719 -1.4307
0.0173 12.6667 950 0.0069 2.2732 -15.6753 0.9900 17.9485 -115.4458 -24.9018 -1.4719 -1.4307
0.0 13.3333 1000 0.0069 2.2757 -15.6836 0.9900 17.9593 -115.4733 -24.8934 -1.4719 -1.4307

Framework versions

  • Transformers 4.41.2
  • Pytorch 2.0.0+cu117
  • Datasets 2.19.2
  • Tokenizers 0.19.1
Downloads last month
5
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tsavage68/UTI_L3_1000steps_1e5rate_03beta_CSFTDPO