8a681bd5e48b1876819df01f82399e28

This model is a fine-tuned version of google/mt5-small on the Helsinki-NLP/opus_books [fi-pl] dataset. It achieves the following results on the evaluation set:

  • Loss: 3.1186
  • Data Size: 1.0
  • Epoch Runtime: 12.7358
  • Bleu: 0.7448

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • total_train_batch_size: 32
  • total_eval_batch_size: 32
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • num_epochs: 50

Training results

Training Loss Epoch Step Validation Loss Data Size Epoch Runtime Bleu
No log 0 0 29.2432 0 1.6355 0.0040
No log 1 70 29.0176 0.0078 2.7695 0.0080
No log 2 140 26.7250 0.0156 1.9340 0.0071
No log 3 210 24.7823 0.0312 2.4169 0.0096
No log 4 280 22.9952 0.0625 2.5776 0.0068
No log 5 350 19.9288 0.125 3.4689 0.0101
No log 6 420 15.9629 0.25 4.9484 0.0116
3.181 7 490 12.7236 0.5 7.0200 0.0051
12.7957 8.0 560 8.4316 1.0 12.1829 0.0183
9.9094 9.0 630 5.9163 1.0 12.0417 0.0204
6.3213 10.0 700 4.2825 1.0 11.0516 0.0636
5.4421 11.0 770 3.8785 1.0 11.8811 0.1150
5.1096 12.0 840 3.7124 1.0 11.1901 0.1584
4.7232 13.0 910 3.6135 1.0 11.4277 0.2136
4.5929 14.0 980 3.5475 1.0 11.3017 0.3249
4.4352 15.0 1050 3.5057 1.0 11.6904 0.3929
4.3344 16.0 1120 3.4632 1.0 11.9123 0.3688
4.3018 17.0 1190 3.4292 1.0 11.8033 0.3894
4.1827 18.0 1260 3.3994 1.0 12.7377 0.4141
4.15 19.0 1330 3.3746 1.0 11.3964 0.4392
4.0598 20.0 1400 3.3483 1.0 11.1041 0.5399
3.9852 21.0 1470 3.3295 1.0 11.9825 0.4914
3.9972 22.0 1540 3.3044 1.0 11.4271 0.5344
3.8871 23.0 1610 3.2907 1.0 11.6221 0.5166
3.881 24.0 1680 3.2773 1.0 12.4179 0.5094
3.8228 25.0 1750 3.2656 1.0 12.6361 0.5314
3.7924 26.0 1820 3.2531 1.0 12.9146 0.5420
3.7602 27.0 1890 3.2372 1.0 11.2274 0.5487
3.7062 28.0 1960 3.2300 1.0 11.4544 0.5856
3.69 29.0 2030 3.2221 1.0 11.5363 0.6032
3.6605 30.0 2100 3.2113 1.0 12.1520 0.6290
3.6333 31.0 2170 3.2026 1.0 12.5515 0.6006
3.6323 32.0 2240 3.1991 1.0 12.3393 0.6086
3.5883 33.0 2310 3.1877 1.0 12.6489 0.6431
3.5566 34.0 2380 3.1791 1.0 12.5892 0.6643
3.5366 35.0 2450 3.1717 1.0 12.8103 0.6764
3.5038 36.0 2520 3.1706 1.0 11.6256 0.7019
3.4822 37.0 2590 3.1654 1.0 11.5844 0.6915
3.4731 38.0 2660 3.1596 1.0 11.6021 0.6762
3.4375 39.0 2730 3.1597 1.0 12.1233 0.6931
3.4068 40.0 2800 3.1518 1.0 12.4407 0.6922
3.3934 41.0 2870 3.1518 1.0 13.1373 0.6849
3.3942 42.0 2940 3.1437 1.0 13.3477 0.7083
3.3409 43.0 3010 3.1393 1.0 13.4745 0.7043
3.3472 44.0 3080 3.1379 1.0 13.1708 0.7050
3.3113 45.0 3150 3.1343 1.0 11.7511 0.7387
3.2914 46.0 3220 3.1290 1.0 12.0681 0.7506
3.2856 47.0 3290 3.1196 1.0 12.2375 0.7463
3.269 48.0 3360 3.1260 1.0 12.4927 0.7377
3.2462 49.0 3430 3.1207 1.0 12.5982 0.7650
3.2483 50.0 3500 3.1186 1.0 12.7358 0.7448

Framework versions

  • Transformers 4.57.0
  • Pytorch 2.8.0+cu128
  • Datasets 4.2.0
  • Tokenizers 0.22.1
Downloads last month
1
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for contemmcm/8a681bd5e48b1876819df01f82399e28

Base model

google/mt5-small
Finetuned
(754)
this model