mt5-large-wmt14-deen

This model is released as part of the work from Are Character-level Translations Worth the Wait? Comparing Character- and Subword-level Models for Machine Translation. It is an mT5 model finetuned on German-->English translation the WMT14 dataset.

To use the model correctly, you must prepend the prompt with "translate X to Y: ", where X and Y are your source and target languages (e.g. German, English).

NOTE: The decoder_start_token_id is 259 for byt5 models and 250099 for mt5 models, which is different from the default token from google's byt5 and mt5 models (which is 0).

Downloads last month: 8

Safetensors

Model size

1B params

Tensor type

F32

Dataset used to train leukas/mt5-large-wmt14-deen

Collection including leukas/mt5-large-wmt14-deen

Are Character-level Translations Worth the Wait?

Collection

Collection of trained models for the paper: Are Character-level Translations Worth the Wait? • 162 items • Updated Sep 11, 2024

Paper for leukas/mt5-large-wmt14-deen

Are Character-level Translations Worth the Wait? Comparing Character- and Subword-level Models for Machine Translation

Paper • 2302.14220 • Published Feb 28, 2023

Evaluation results

BLEU on wmt14
test set verified

15.919
loss on wmt14
test set verified

1.098
gen_len on wmt14
test set verified

19.087