Instructions to use mainlp/copy-first-translate-later with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mainlp/copy-first-translate-later with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mainlp/copy-first-translate-later")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("mainlp/copy-first-translate-later") model = AutoModelForCausalLM.from_pretrained("mainlp/copy-first-translate-later", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mainlp/copy-first-translate-later with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mainlp/copy-first-translate-later" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mainlp/copy-first-translate-later", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/mainlp/copy-first-translate-later
- SGLang
How to use mainlp/copy-first-translate-later with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mainlp/copy-first-translate-later" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mainlp/copy-first-translate-later", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mainlp/copy-first-translate-later" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mainlp/copy-first-translate-later", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use mainlp/copy-first-translate-later with Docker Model Runner:
docker model run hf.co/mainlp/copy-first-translate-later
copy-first-translate-later
These are the model checkpoints trained and studied in Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining.
A 1.67B-parameter Llama-style language model pretrained from scratch on 9 languages. We release all 200 intermediate checkpoints (every 185M tokens/500 steps, step 500 to step 100,000), together with the full training state and a record of exactly which training documents were seen between consecutive checkpoints.
Checkpoints
Each checkpoint is a separate branch named step{N} (e.g. step500,
step1000, …, step100000). Every branch contains:
| file | description |
|---|---|
model.safetensors |
model weights (fp32, standard LlamaForCausalLM parameter names) |
config.json, tokenizer*.json, special_tokens_map.json |
model config (as used in training) and tokenizer |
optimizer.bin, scheduler.bin, random_states_*.pkl |
full 🤗 Accelerate training state (8 ranks), for resuming training |
training_data_seen.parquet |
the documents seen since the previous checkpoint (see Training data) |
Loading a checkpoint
Always pass a revision: main holds only the config and tokenizer, not
weights.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "mainlp/copy-first-translate-later"
REVISION = "step100000"
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REVISION, dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained(REPO, revision=REVISION)
The tokenizer is the unmodified Mistral-Nemo tokenizer
(mistralai/Mistral-Nemo-Base-2407), with pad_token set to eos_token as
in training. The config uses dynamic RoPE scaling with factor 1.0, as in
training: this has no effect up to the 4,096-token training context, but
changes the RoPE base beyond it.
Training details
| architecture | Llama; 24 layers, hidden size 2048, MLP size 5632, 16 attention heads / 8 KV heads, RoPE θ = 500,000, untied embeddings |
| parameters | 1.67B |
| tokenizer | mistralai/Mistral-Nemo-Base-2407 (vocab 131,072) |
| training steps | 100,000 |
| batch size | 512 documents per step (8 GPUs × 2 × 32 gradient-accumulation steps); the 2 documents of each micro-batch are packed into one sequence, each truncated to 4,096 tokens |
| optimizer | AdamW, lr 1e-4, β = (0.9, 0.999), ε = 1e-8, weight decay 0.01 |
| schedule | warmup-stable-decay: 10% linear warmup, 10% 1-sqrt decay to 0.1 × peak lr |
| precision | bf16 mixed precision |
| seed | 42 |
Training data
The model was trained on unmodified documents from two public datasets:
- English (
eng_Latn): FineWeb-HQ - Chinese, German, Spanish, Japanese, Italian, Portuguese, Indonesian,
French (
cmn_Hani,deu_Latn,spa_Latn,jpn_Jpan,ita_Latn,por_Latn,ind_Latn,fra_Latn): FineWeb2-HQ
Languages were sampled with probability 0.5 for English and 0.0625 for each of the other eight.
Which documents were seen when
training_data_seen.parquet on branch step{N} lists every document the
training data loader consumed between checkpoint step{N-500} and
checkpoint step{N} (for step500: since the start of training). Every
checkpoint interval covers 256,000 documents.
| column | description |
|---|---|
id |
the document's FineWeb / FineWeb2 id (e.g. <urn:uuid:…>) |
url |
source URL |
dump |
Common Crawl snapshot (e.g. CC-MAIN-2016-50) |
language |
language subset, matching the FineWeb2-HQ config names (eng_Latn for English) |
step |
optimizer step whose update included the document (N-499 … N) |
micro_batch |
gradient-accumulation micro-batch within the step (0–31) |
rank |
data-parallel rank (GPU) that processed the document (0–7) |
slot |
position within the micro-batch (0 or 1); both documents are packed into one sequence |
Notes:
- Rows are in training order, sorted by (
step,micro_batch,rank,slot). Micro-batches with the samestepandmicro_batchwere processed in parallel on all 8 ranks, so there is no order between ranks. - Document text is not included. Join on
idagainst FineWeb-HQ / FineWeb2-HQ (or the full FineWeb / FineWeb-2, which use the same IDs) to recover it. FineWeb-HQ is organised bydump, so you only need to download the snapshots that appear in the manifest. - All data seen up to step N is the union of the manifests of all branches
step500…step{N}. - No document was seen twice through the data loader, but two
ids occur twice across all manifests because the same document appears in two source files:<urn:uuid:7e738bbd-163c-473b-b159-58dc10fa9761>(twice inind_Latn) and<urn:uuid:fbd75fee-d88d-4105-bd02-556915ac3c09>(once ineng_Latn, once inind_Latn).
import pandas as pd
from huggingface_hub import hf_hub_download
seen = pd.read_parquet(
hf_hub_download(
"mainlp/copy-first-translate-later",
"training_data_seen.parquet",
revision="step1000",
)
)
print(seen["language"].value_counts())
Data licensing and attribution
FineWeb-HQ and FineWeb2-HQ (and the FineWeb and FineWeb-2 datasets they are drawn from) are released under the Open Data Commons Attribution License (ODC-By) v1.0, and their use is subject to Common Crawl's Terms of Use. The per-checkpoint manifests are derived from these datasets and are distributed under the same terms.
Removal requests
The manifests list source URLs. To have content removed from the upstream datasets, use the FineWeb opt-out form.
Citation
If you use these models please cite the following (to be published at EMNLP 2026).
@misc{koerner2026copyfirsttranslatelater,
title={Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining},
author={Felicia Körner and Maria Matveev and Florian Eichin and Gitta Kutyniok and Barbara Plank and Michael A. Hedderich},
year={2026},
eprint={2604.17633},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.17633}
}
- Downloads last month
- 413