Hyponatremia_L3_1000steps_1e6rate_01beta_CSFTDPO

This model is a fine-tuned version of tsavage68/Summary4500_L3_100steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0014
  • Rewards/chosen: -1.4084
  • Rewards/rejected: -18.4001
  • Rewards/accuracies: 0.9980
  • Rewards/margins: 16.9917
  • Logps/rejected: -317.1989
  • Logps/chosen: -98.2741
  • Logits/rejected: -1.0846
  • Logits/chosen: -1.0076

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-06
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss Rewards/chosen Rewards/rejected Rewards/accuracies Rewards/margins Logps/rejected Logps/chosen Logits/rejected Logits/chosen
0.0114 0.0112 50 0.0093 -0.2009 -5.9855 0.9980 5.7846 -193.0523 -86.1985 -1.1075 -1.0649
0.0 0.0224 100 0.0024 -0.9378 -10.4848 0.9980 9.5470 -238.0455 -93.5676 -1.1001 -1.0461
0.0 0.0336 150 0.0017 -1.0803 -12.7703 0.9980 11.6899 -260.8999 -94.9929 -1.0979 -1.0362
0.0 0.0448 200 0.0015 -2.1051 -16.1714 0.9980 14.0663 -294.9110 -105.2404 -1.0968 -1.0306
0.0 0.0559 250 0.0015 -1.2418 -15.6144 0.9980 14.3726 -289.3413 -96.6073 -1.0946 -1.0268
0.0 0.0671 300 0.0015 -1.2850 -16.0588 0.9980 14.7738 -293.7853 -97.0396 -1.0920 -1.0240
0.0 0.0783 350 0.0014 -1.5607 -17.5217 0.9980 15.9609 -308.4142 -99.7972 -1.0919 -1.0200
0.0 0.0895 400 0.0014 -1.5463 -17.5816 0.9980 16.0353 -309.0129 -99.6524 -1.0908 -1.0187
0.0 0.1007 450 0.0014 -1.5768 -17.6781 0.9980 16.1012 -309.9779 -99.9583 -1.0908 -1.0182
0.0 0.1119 500 0.0014 -1.4380 -17.9331 0.9980 16.4952 -312.5286 -98.5695 -1.0817 -1.0071
0.0 0.1231 550 0.0014 -1.4831 -18.1851 0.9980 16.7020 -315.0485 -99.0211 -1.0852 -1.0099
0.0 0.1343 600 0.0014 -1.4779 -18.1900 0.9980 16.7121 -315.0977 -98.9690 -1.0853 -1.0100
0.0 0.1454 650 0.0014 -1.4375 -18.2718 0.9980 16.8342 -315.9149 -98.5652 -1.0861 -1.0096
0.0 0.1566 700 0.0014 -1.4049 -18.3712 0.9980 16.9664 -316.9096 -98.2383 -1.0854 -1.0084
0.0004 0.1678 750 0.0014 -1.4073 -18.3876 0.9980 16.9803 -317.0729 -98.2626 -1.0845 -1.0075
0.0 0.1790 800 0.0014 -1.4175 -18.4190 0.9980 17.0016 -317.3878 -98.3644 -1.0846 -1.0076
0.0001 0.1902 850 0.0014 -1.4088 -18.4040 0.9980 16.9952 -317.2370 -98.2774 -1.0844 -1.0074
0.0 0.2014 900 0.0014 -1.4115 -18.4067 0.9980 16.9952 -317.2642 -98.3050 -1.0845 -1.0074
0.0 0.2126 950 0.0014 -1.4069 -18.4091 0.9980 17.0022 -317.2884 -98.2590 -1.0845 -1.0075
0.0 0.2238 1000 0.0014 -1.4084 -18.4001 0.9980 16.9917 -317.1989 -98.2741 -1.0846 -1.0076

Framework versions

  • Transformers 4.42.4
  • Pytorch 2.0.0+cu117
  • Datasets 2.20.0
  • Tokenizers 0.19.1
Downloads last month
8
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tsavage68/Summary4500_L3_1000steps_1e6rate_01beta_CSFTDPO