58fceacaa752d5e3474d761054b94d40

This model is a fine-tuned version of facebook/mbart-large-cc25 on the Helsinki-NLP/opus_books [fr-nl] dataset. It achieves the following results on the evaluation set:

  • Loss: 2.0810
  • Data Size: 1.0
  • Epoch Runtime: 264.9675
  • Bleu: 10.1440

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • total_train_batch_size: 32
  • total_eval_batch_size: 32
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • num_epochs: 50

Training results

Training Loss Epoch Step Validation Loss Data Size Epoch Runtime Bleu
No log 0 0 10.7814 0 23.0063 0.1624
No log 1 1000 3.7832 0.0078 24.3671 2.8235
No log 2 2000 3.4683 0.0156 26.5974 3.6917
No log 3 3000 2.9326 0.0312 31.0007 4.4876
0.1208 4 4000 2.5241 0.0625 39.1177 5.4169
2.5342 5 5000 2.2655 0.125 54.3441 6.1478
0.1364 6 6000 2.0353 0.25 83.4401 7.2585
0.171 7 7000 1.8527 0.5 142.1893 8.4845
1.657 8.0 8000 1.7065 1.0 262.8612 16.5878
1.424 9.0 9000 1.6926 1.0 259.6856 13.0326
1.2043 10.0 10000 1.6815 1.0 261.0464 10.3316
1.0234 11.0 11000 1.7398 1.0 260.1502 10.6781
0.8361 12.0 12000 1.8461 1.0 260.9129 11.2870
0.6943 13.0 13000 1.9749 1.0 260.6425 11.4679
0.5505 14.0 14000 2.0810 1.0 264.9675 10.1440

Framework versions

  • Transformers 4.57.0
  • Pytorch 2.8.0+cu128
  • Datasets 4.2.0
  • Tokenizers 0.22.1
Downloads last month
1
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for contemmcm/58fceacaa752d5e3474d761054b94d40

Finetuned
(68)
this model