12000samples

This model is a fine-tuned version of Helsinki-NLP/opus-mt-es-es on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 0.1451
  • Bleu Msl: 92.1051
  • Bleu Asl: 0
  • Ter Msl: 4.7161
  • Ter Asl: 100

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 32
  • eval_batch_size: 64
  • seed: 42
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 30
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Bleu Msl Bleu Asl Ter Msl Ter Asl
No log 1.0 386 0.9869 20.9686 82.0457 66.6038 9.1959
1.5187 2.0 772 0.4370 48.4639 86.5183 36.0377 6.8058
0.5406 3.0 1158 0.2964 60.2032 88.7852 26.3208 5.6917
0.2888 4.0 1544 0.2504 65.3635 90.1910 22.3585 4.9828
0.2888 5.0 1930 0.2023 72.9532 91.6531 17.4528 4.2941
0.2019 6.0 2316 0.1680 62.1974 92.4105 19.8113 3.9093
0.1391 7.0 2702 0.1532 75.8711 92.9684 15.2830 3.6864
0.1058 8.0 3088 0.1407 47.4649 93.0796 26.0377 3.5852
0.1058 9.0 3474 0.1361 77.0535 93.1579 14.2453 3.5042
0.0868 10.0 3860 0.1288 59.0665 93.5742 18.4906 3.3421
0.0729 11.0 4246 0.1239 80.7391 93.4642 12.6415 3.3826
0.0629 12.0 4632 0.1206 49.3562 93.5802 23.6792 3.2408
0.054 13.0 5018 0.1185 77.7535 94.0327 12.8302 3.0788
0.054 14.0 5404 0.1170 79.9193 93.7494 12.4528 3.1193
0.0481 15.0 5790 0.1133 79.8567 93.7407 11.4151 2.9978
0.0425 16.0 6176 0.1131 83.0857 93.6535 10.7547 2.8560
0.0385 17.0 6562 0.1121 82.8013 94.0053 10.7547 2.6737
0.0385 18.0 6948 0.1112 84.5147 93.9244 10.1887 2.6939
0.0359 19.0 7334 0.1108 83.7256 94.1014 10.0943 2.6939
0.032 20.0 7720 0.1103 82.0662 94.2459 10.4717 2.7952
0.0298 21.0 8106 0.1099 84.3243 94.5135 10.1887 2.6939
0.0298 22.0 8492 0.1105 83.9736 94.5280 10.1887 2.7345
0.0275 23.0 8878 0.1091 84.6834 94.4765 10.0 2.7547
0.0261 24.0 9264 0.1100 84.3382 94.6091 10.1887 2.6737
0.0247 25.0 9650 0.1099 83.8395 94.5048 10.0943 2.7750
0.024 26.0 10036 0.1104 83.1135 94.5483 10.1887 2.7547
0.024 27.0 10422 0.1100 83.5175 94.5483 10.2830 2.7547
0.0227 28.0 10808 0.1101 84.2731 94.5440 10.0943 2.7750
0.022 29.0 11194 0.1100 84.3292 94.5440 10.0943 2.7750
0.0224 30.0 11580 0.1101 83.8359 94.5440 10.1887 2.7750

Framework versions

  • Transformers 4.46.2
  • Pytorch 2.5.1+cu121
  • Datasets 3.1.0
  • Tokenizers 0.20.3
Downloads last month
9
Safetensors
Model size
61.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vania2911/12000samples

Finetuned
(57)
this model