Instructions to use cadazar/han2han-pt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cadazar/han2han-pt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cadazar/han2han-pt", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("cadazar/han2han-pt", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Han2Han PT
The pre-trained (PT) checkpoint of Han2Han, a 169M-parameter encoder-decoder model that learns script-invariant representations of Korean text: a document written in Hanja and its Hangul transcription land at the same point in embedding space. The recipe (jamo and character-level embedding fusion, morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described in the paper, accepted to Findings of EMNLP 2026.
This is the starting point for fine-tuning, and the first of three checkpoints:
| Repo | Stage |
|---|---|
cadazar/han2han-pt (this one) |
pre-training, 35B tokens |
cadazar/han2han-it |
instruction tuning from these weights |
cadazar/han2han-rl |
reinforcement learning on transcription in both directions (Hangul to Hanja and Hanja to Hangul), from han2han-it |
These weights are an average
The weights are the uniform average of the last five checkpoints of the
pre-training run (steps 6347, 6441, 6536, 6630 and 6725, saved every 0.5B
tokens from 33B to 35B), not the final checkpoint alone. The average is what
the instruction tuning started from, chosen after comparing starting points,
so it is the pre-trained model that han2han-it and han2han-rl descend
from. The learning rate had nearly finished its cooldown over those steps: no
weight differs from the final checkpoint by more than 0.0015.
Intended use
Fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike.
It is not a chat model. The tokenizer and the chat template are the same as in
the other two repos, but pre-training never used the chat tokens
(<|system|>, <|user|>, <|assistant|>, <|think|>, <|end_of_turn|>):
their embedding rows are still nearly one shared vector here, as for every
other token pre-training never saw. For generation from a prompt use
han2han-it or han2han-rl.
Usage
Runtime requirements: torch and transformers (tested with torch 2.14.1 and
transformers 5.18.0 on CPU). The model class ships in this repo, so loading needs
trust_remote_code=True; the tokenizer itself runs no custom code, but without
the flag AutoTokenizer stops to ask.
output_sentence_embeddings=True returns the mean-pooled encoder states, one
640-dimensional vector per input.
import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
repo = "cadazar/han2han-pt"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
texts = [
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
"ํ์ฅ์ ์ผ์ํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.",
"ๅ็ต๋ ๅๅๅญ์์๋ ็ ดๅขจ์ ๅฆๆณ์ผ๋ก ็ฝ้ช์ ่ฑกๅพตํ ์ ์๋ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9315, 0.8253],
# [0.9315, 1.0000, 0.7910],
# [0.8253, 0.7910, 1.0000]])
The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.
Tokenizer
tokenizer.json is a tokenizers-library build of the SentencePiece model
(spiece.model, kept here as the source), made by
scripts/build_hf_tokenizer.py in the GitHub repo. Calling the tokenizer does
not add BOS or EOS tokens, and special tokens written into the text are mapped
to their ids. It encodes like the SentencePiece wrapper the training code uses,
with one known difference: where two segmentations of a span have exactly the
same score (runs of digits, mostly), the same pieces can come out in a different
order.
Files
| File | Contents |
|---|---|
model.safetensors |
fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables |
config.json, generation_config.json |
model and generation config, with auto_map entries for the Auto classes |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
fast tokenizer (38400 pieces) and chat template, the same as in han2han-it |
spiece.model |
the SentencePiece model tokenizer.json was built from |
modeling_han2han.py, han2han_config.py |
modeling code from the GitHub repo at commit af1330e, the same files as in han2han-it |
Citation
@inproceedings{han2han2026,
title = {Han2Han: Efficient Language-Specific Character Representation
through Script-Aware Pre-Training for Historical Text Analysis},
author = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}
License
Apache License 2.0, the same as the GitHub repo.
- Downloads last month
- 31