Hyponatremia_L3_1000steps_1e6rate_05beta_CSFTDPO

This model is a fine-tuned version of tsavage68/Summary4500_L3_100steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0014
  • Rewards/chosen: 1.0176
  • Rewards/rejected: -20.0926
  • Rewards/accuracies: 0.9980
  • Rewards/margins: 21.1103
  • Logps/rejected: -173.3826
  • Logps/chosen: -82.1545
  • Logits/rejected: -1.0918
  • Logits/chosen: -1.0524

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-06
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss Rewards/chosen Rewards/rejected Rewards/accuracies Rewards/margins Logps/rejected Logps/chosen Logits/rejected Logits/chosen
0.0003 0.0112 50 0.0023 0.5359 -8.1120 0.9980 8.6479 -149.4213 -83.1180 -1.1012 -1.0676
0.0 0.0224 100 0.0016 0.1709 -10.9251 0.9980 11.0960 -155.0475 -83.8480 -1.1026 -1.0673
0.0 0.0336 150 0.0014 -0.1278 -13.7945 0.9980 13.6667 -160.7863 -84.4453 -1.1022 -1.0664
0.0 0.0448 200 0.0014 -0.0574 -14.6683 0.9980 14.6109 -162.5339 -84.3046 -1.1016 -1.0657
0.0 0.0559 250 0.0014 0.3311 -15.4389 0.9980 15.7700 -164.0751 -83.5275 -1.0992 -1.0628
0.0 0.0671 300 0.0014 0.3433 -15.4472 0.9980 15.7905 -164.0917 -83.5031 -1.0990 -1.0626
0.0 0.0783 350 0.0014 0.4029 -17.0508 0.9980 17.4537 -167.2989 -83.3839 -1.1027 -1.0639
0.0 0.0895 400 0.0014 0.3792 -17.1575 0.9980 17.5367 -167.5124 -83.4315 -1.1026 -1.0637
0.0 0.1007 450 0.0014 0.4159 -17.1507 0.9980 17.5667 -167.4988 -83.3579 -1.1033 -1.0647
0.0 0.1119 500 0.0014 0.6555 -18.5577 0.9980 19.2132 -170.3127 -82.8788 -1.0977 -1.0583
0.0 0.1231 550 0.0014 0.9891 -20.0773 0.9980 21.0664 -173.3519 -82.2115 -1.0934 -1.0539
0.0 0.1343 600 0.0014 0.9858 -20.0819 0.9980 21.0676 -173.3611 -82.2182 -1.0935 -1.0539
0.0 0.1454 650 0.0014 0.9858 -20.0819 0.9980 21.0676 -173.3611 -82.2182 -1.0935 -1.0539
0.0 0.1566 700 0.0014 0.9752 -20.1001 0.9980 21.0753 -173.3975 -82.2393 -1.0933 -1.0536
0.0 0.1678 750 0.0014 0.9974 -20.1078 0.9980 21.1052 -173.4129 -82.1949 -1.0923 -1.0527
0.0 0.1790 800 0.0014 1.0079 -20.1039 0.9980 21.1118 -173.4052 -82.1740 -1.0923 -1.0528
0.0 0.1902 850 0.0014 1.0134 -20.1134 0.9980 21.1268 -173.4241 -82.1630 -1.0920 -1.0524
0.0 0.2014 900 0.0014 1.0201 -20.0711 0.9980 21.0912 -173.3395 -82.1496 -1.0918 -1.0524
0.0 0.2126 950 0.0014 1.0208 -20.0898 0.9980 21.1107 -173.3770 -82.1481 -1.0918 -1.0524
0.0 0.2238 1000 0.0014 1.0176 -20.0926 0.9980 21.1103 -173.3826 -82.1545 -1.0918 -1.0524

Framework versions

  • Transformers 4.42.4
  • Pytorch 2.0.0+cu117
  • Datasets 2.20.0
  • Tokenizers 0.19.1
Downloads last month
6
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tsavage68/Summary4500_L3_1000steps_1e6rate_05beta_CSFTDPO