Text Generation
Transformers
llama

copy-first-translate-later

These are the model checkpoints trained and studied in Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining.

A 1.67B-parameter Llama-style language model pretrained from scratch on 9 languages. We release all 200 intermediate checkpoints (every 185M tokens/500 steps, step 500 to step 100,000), together with the full training state and a record of exactly which training documents were seen between consecutive checkpoints.

Checkpoints

Each checkpoint is a separate branch named step{N} (e.g. step500, step1000, …, step100000). Every branch contains:

file description
model.safetensors model weights (fp32, standard LlamaForCausalLM parameter names)
config.json, tokenizer*.json, special_tokens_map.json model config (as used in training) and tokenizer
optimizer.bin, scheduler.bin, random_states_*.pkl full 🤗 Accelerate training state (8 ranks), for resuming training
training_data_seen.parquet the documents seen since the previous checkpoint (see Training data)

Loading a checkpoint

Always pass a revision: main holds only the config and tokenizer, not weights.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "mainlp/copy-first-translate-later"
REVISION = "step100000"

model = AutoModelForCausalLM.from_pretrained(REPO, revision=REVISION, dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained(REPO, revision=REVISION)

The tokenizer is the unmodified Mistral-Nemo tokenizer (mistralai/Mistral-Nemo-Base-2407), with pad_token set to eos_token as in training. The config uses dynamic RoPE scaling with factor 1.0, as in training: this has no effect up to the 4,096-token training context, but changes the RoPE base beyond it.

Training details

architecture Llama; 24 layers, hidden size 2048, MLP size 5632, 16 attention heads / 8 KV heads, RoPE θ = 500,000, untied embeddings
parameters 1.67B
tokenizer mistralai/Mistral-Nemo-Base-2407 (vocab 131,072)
training steps 100,000
batch size 512 documents per step (8 GPUs × 2 × 32 gradient-accumulation steps); the 2 documents of each micro-batch are packed into one sequence, each truncated to 4,096 tokens
optimizer AdamW, lr 1e-4, β = (0.9, 0.999), ε = 1e-8, weight decay 0.01
schedule warmup-stable-decay: 10% linear warmup, 10% 1-sqrt decay to 0.1 × peak lr
precision bf16 mixed precision
seed 42

Training data

The model was trained on unmodified documents from two public datasets:

  • English (eng_Latn): FineWeb-HQ
  • Chinese, German, Spanish, Japanese, Italian, Portuguese, Indonesian, French (cmn_Hani, deu_Latn, spa_Latn, jpn_Jpan, ita_Latn, por_Latn, ind_Latn, fra_Latn): FineWeb2-HQ

Languages were sampled with probability 0.5 for English and 0.0625 for each of the other eight.

Which documents were seen when

training_data_seen.parquet on branch step{N} lists every document the training data loader consumed between checkpoint step{N-500} and checkpoint step{N} (for step500: since the start of training). Every checkpoint interval covers 256,000 documents.

column description
id the document's FineWeb / FineWeb2 id (e.g. <urn:uuid:…>)
url source URL
dump Common Crawl snapshot (e.g. CC-MAIN-2016-50)
language language subset, matching the FineWeb2-HQ config names (eng_Latn for English)
step optimizer step whose update included the document (N-499 … N)
micro_batch gradient-accumulation micro-batch within the step (0–31)
rank data-parallel rank (GPU) that processed the document (0–7)
slot position within the micro-batch (0 or 1); both documents are packed into one sequence

Notes:

  • Rows are in training order, sorted by (step, micro_batch, rank, slot). Micro-batches with the same step and micro_batch were processed in parallel on all 8 ranks, so there is no order between ranks.
  • Document text is not included. Join on id against FineWeb-HQ / FineWeb2-HQ (or the full FineWeb / FineWeb-2, which use the same IDs) to recover it. FineWeb-HQ is organised by dump, so you only need to download the snapshots that appear in the manifest.
  • All data seen up to step N is the union of the manifests of all branches step500 … step{N}.
  • No document was seen twice through the data loader, but two ids occur twice across all manifests because the same document appears in two source files: <urn:uuid:7e738bbd-163c-473b-b159-58dc10fa9761> (twice in ind_Latn) and <urn:uuid:fbd75fee-d88d-4105-bd02-556915ac3c09> (once in eng_Latn, once in ind_Latn).
import pandas as pd
from huggingface_hub import hf_hub_download

seen = pd.read_parquet(
    hf_hub_download(
        "mainlp/copy-first-translate-later",
        "training_data_seen.parquet",
        revision="step1000",
    )
)
print(seen["language"].value_counts())

Data licensing and attribution

FineWeb-HQ and FineWeb2-HQ (and the FineWeb and FineWeb-2 datasets they are drawn from) are released under the Open Data Commons Attribution License (ODC-By) v1.0, and their use is subject to Common Crawl's Terms of Use. The per-checkpoint manifests are derived from these datasets and are distributed under the same terms.

Removal requests

The manifests list source URLs. To have content removed from the upstream datasets, use the FineWeb opt-out form.

Citation

If you use these models please cite the following (to be published at EMNLP 2026).

@misc{koerner2026copyfirsttranslatelater,
      title={Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining},
      author={Felicia Körner and Maria Matveev and Florian Eichin and Gitta Kutyniok and Barbara Plank and Michael A. Hedderich},
      year={2026},
      eprint={2604.17633},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2604.17633}
}
Downloads last month
413
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train mainlp/copy-first-translate-later

Paper for mainlp/copy-first-translate-later