Hyponatremia_L3_1000steps_1e6rate_03beta_CSFTDPO

This model is a fine-tuned version of tsavage68/Summary4500_L3_100steps_1e6rate_SFT on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0014
  • Rewards/chosen: -0.1633
  • Rewards/rejected: -19.7875
  • Rewards/accuracies: 0.9980
  • Rewards/margins: 19.6243
  • Logps/rejected: -199.1558
  • Logps/chosen: -84.7341
  • Logits/rejected: -1.0916
  • Logits/chosen: -1.0431

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-06
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_steps: 100
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss Rewards/chosen Rewards/rejected Rewards/accuracies Rewards/margins Logps/rejected Logps/chosen Logits/rejected Logits/chosen
0.0011 0.0112 50 0.0032 0.4589 -7.1323 0.9980 7.5912 -156.9717 -82.6602 -1.1014 -1.0652
0.0 0.0224 100 0.0017 0.0259 -10.1984 0.9980 10.2243 -167.1920 -84.1034 -1.1010 -1.0621
0.0 0.0336 150 0.0015 -0.2730 -12.2233 0.9980 11.9503 -173.9416 -85.0998 -1.1007 -1.0606
0.0 0.0448 200 0.0014 -0.2383 -14.0974 0.9980 13.8592 -180.1888 -84.9840 -1.0957 -1.0547
0.0 0.0559 250 0.0014 -0.4961 -16.6298 0.9980 16.1337 -188.6300 -85.8433 -1.0906 -1.0485
0.0 0.0671 300 0.0014 -0.4855 -16.6491 0.9980 16.1636 -188.6945 -85.8082 -1.0906 -1.0484
0.0 0.0783 350 0.0014 -0.4651 -18.0207 0.9980 17.5556 -193.2663 -85.7401 -1.0930 -1.0475
0.0 0.0895 400 0.0014 -0.4705 -18.0770 0.9980 17.6065 -193.4542 -85.7582 -1.0925 -1.0469
0.0 0.1007 450 0.0014 -0.4749 -18.1128 0.9980 17.6379 -193.5734 -85.7727 -1.0927 -1.0470
0.0 0.1119 500 0.0014 -0.4497 -18.3137 0.9980 17.8641 -194.2431 -85.6886 -1.0920 -1.0462
0.0 0.1231 550 0.0014 -0.1952 -19.8131 0.9980 19.6179 -199.2410 -84.8404 -1.0929 -1.0442
0.0 0.1343 600 0.0014 -0.1956 -19.8283 0.9980 19.6327 -199.2916 -84.8418 -1.0929 -1.0442
0.0 0.1454 650 0.0014 -0.1887 -19.8240 0.9980 19.6353 -199.2772 -84.8187 -1.0930 -1.0444
0.0 0.1566 700 0.0014 -0.1862 -19.8230 0.9980 19.6368 -199.2740 -84.8106 -1.0930 -1.0443
0.0 0.1678 750 0.0014 -0.1676 -19.7855 0.9980 19.6180 -199.1491 -84.7483 -1.0918 -1.0432
0.0 0.1790 800 0.0014 -0.1614 -19.7862 0.9980 19.6248 -199.1514 -84.7279 -1.0917 -1.0430
0.0 0.1902 850 0.0014 -0.1737 -19.8108 0.9980 19.6371 -199.2332 -84.7688 -1.0916 -1.0433
0.0 0.2014 900 0.0014 -0.1638 -19.8003 0.9980 19.6364 -199.1983 -84.7359 -1.0916 -1.0432
0.0 0.2126 950 0.0014 -0.1645 -19.7862 0.9980 19.6217 -199.1513 -84.7380 -1.0916 -1.0431
0.0 0.2238 1000 0.0014 -0.1633 -19.7875 0.9980 19.6243 -199.1558 -84.7341 -1.0916 -1.0431

Framework versions

  • Transformers 4.42.4
  • Pytorch 2.0.0+cu117
  • Datasets 2.20.0
  • Tokenizers 0.19.1
Downloads last month
6
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tsavage68/Summary4500_L3_1000steps_1e6rate_03beta_CSFTDPO