Instructions to use cadazar/han2han-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cadazar/han2han-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cadazar/han2han-it", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("cadazar/han2han-it", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Han2Han IT
Instruction-tuned checkpoint of Han2Han, a 169M-parameter encoder-decoder model that learns script-invariant representations of Korean text: a document written in Hanja and its Hangul transcription land at the same point in embedding space. The recipe (jamo and character-level embedding fusion, morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described in the paper, accepted to Findings of EMNLP 2026.
Two versions live in this repo.
| Revision | Weights | Use it for |
|---|---|---|
main |
the instruction tuning continued for 30M tokens in fp32 (step 762 of the continuation) | generation and chat: it ends its turns |
emnlp2026 |
han2han-ul2-base-1-it, step 43153, the checkpoint behind the paper |
reproducing the paper |
The first checkpoint was trained with its weights in bf16. The updates to the
embedding rows of the chat tokens were too small to survive bf16 rounding, so
those rows barely moved from where pre-training left them, <|end_of_turn|>
among them, and the model did not learn to end a turn: sampled at temperature 0.7, none of
its outputs stopped. main continues the same instruction tuning (the same
four task families, learning rate 3e-5) in fp32 until those rows are trained.
Nothing else about the recipe changed. To load the paper's weights, pass
revision="emnlp2026" to from_pretrained.
On 256 newspaper sentences from the 1920s, greedy decoding with the task prompt in the system message:
emnlp2026 |
main |
|
|---|---|---|
| Hangul to Hanja: share of the original Hanja restored | 0.19 | 0.31 |
| Hanja to Hangul: share of the Hanja read correctly | 0.65 | 0.89 |
| outputs that end on their own (to Hanja / to Hangul) | 76% / 91% | 99% / 100% |
| the same, sampled at temperature 0.7 | 0% / 0% | 100% / 100% |
| three articles in slices of 250 to 500 characters: Hanja restored | 0.66 | 0.81 |
The sentences are from the validation and test splits of the sentence-level transcription set. The instruction tuning used the article-level set and the same newspapers were part of pre-training, so read the table as a comparison of the two checkpoints, not as performance on unseen text.
The sentence embeddings of main have not been run through the paper's
evaluations; for those results use emnlp2026.
cadazar/han2han-pt is the
pre-trained checkpoint both started from, and
cadazar/han2han-rl is main
trained further on transcription with reinforcement learning, which is the one
to use for converting between the scripts.
This repo holds the PyTorch weights, the tokenizer with its chat template, and
the modeling code needed to load them through the transformers Auto classes
with trust_remote_code=True. Training code, the Flax model, the Flax-to-PyTorch
converter, and the evaluation pipeline live in the GitHub repo.
Intended use
Han2Han is meant for representations: classification and fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike. The model generates text, but generation is not what it was built or evaluated for, and improving it is follow-up work:
- Prompt layout. The task prompt goes in the system message and the text alone in the user message, which is how the instruction tuning presented every task. With the instruction written into the user message instead, Hangul to Hanja restoration returns the Hangul input unchanged.
- Hangul to Hanja restoration (
ํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:) works on article-length input. Given the first 478 characters of a 1922 newspaper article in Hangul, greedy decoding restores 84% of its 251 Hanja. The article comes from a corpus used in pre-training, so this is an illustration and not a held-out measurement. On single sentences it restores much less (0.31 above), and a third of the outputs come back with no Hanja at all: the instruction tuning had articles for this task and no single sentences. - Hanja to Hangul transcription (
ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:) works on sentences and articles, but can leave some Hanja untranscribed or change a word. - Layout matters in both directions. With the instruction in the user message, Hangul to Hanja restores under 1% of the Hanja and Hanja to Hangul reads 45% of them.
- Stopping. The outputs measured here end on their own, in fp32. Give
generate()a token limit anyway: twice the input's tokens plus 8 leaves room for a correct answer and cuts off the occasional output that repeats itself (about 4% of the sentences).
hanja_transcription_demo.ipynb in the GitHub repo was written with the other
layout (the instruction in the user message), so the Hangul to Hanja outputs in
its sections 4 to 9 understate what this checkpoint does. Its section 10 runs
the layout described here.
Usage
Runtime requirements: torch and transformers (tested with torch 2.14.1 and
transformers 5.18.0 on CPU). The model class ships in this repo, so loading needs
trust_remote_code=True; the tokenizer itself runs no custom code, but without
the flag AutoTokenizer stops to ask.
Sentence embeddings
output_sentence_embeddings=True returns the mean-pooled encoder states, one
640-dimensional vector per input.
import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
texts = [
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
"ํ์ฅ์ ์ผ์ํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.",
"ๅ็ต๋ ๅๅๅญ์์๋ ็ ดๅขจ์ ๅฆๆณ์ผ๋ก ็ฝ้ช์ ่ฑกๅพตํ ์ ์๋ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9020, 0.8546],
# [0.9020, 1.0000, 0.8100],
# [0.8546, 0.8100, 1.0000]])
The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.
Chat
AutoModelForCausalLM loads Han2HanForCausalLM, a wrapper that gives the
encoder-decoder the interface chat tooling expects: generate() takes the
rendered chat prompt as one sequence and returns the prompt followed by the
reply. pipeline("text-generation") works with it, and so does the
transformers server:
transformers serve --trust-remote-code
transformers chat cadazar/han2han-it --system-prompt "ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:"
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()
def run(task_prompt, text):
messages = [
{"role": "system", "content": task_prompt},
{"role": "user", "content": text},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
)
budget = 2 * len(tokenizer(text).input_ids) + 8
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=budget, do_sample=False)
return tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(run(
"ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:",
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
))
# ํ์์ฅ์ ์ผ์ํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.
hangul = (
"๋ฏธ์ ์ ๋ ๊ณํ ์ด๋
๋ถ ์ฌ์์ผ๋ก ๋ช
๋
์ด์ฐจ ๊ฐ์ค์ด๋
๋ถ์์๋ ์กฐ์ ์ ์ฌ ํ ๋ฏธ์ ์ ๋ฐ๋ฌ์ ๋น๋ณดํ "
"๋ชฉ์ ์ผ๋ก ๋๊ฒฝ์ ์ ๊ตญ๋ฏธ์ ์์ ๋ํ๋ฅผ ๋ฐฉํ ์ฌ ๋งค๋
์ผ์ฐจ ๋ฏธ์ ์ ๋ํ๋ฅผ ๊ฐํ ๋ฐฉ์นจ์ ๋ด์ ํ๊ณ ์ "
"์ด์ญ์ก์ผ ์ค์ ์ญ์ ์ด๋
๋ถ ์ ์ดํ์์ค์ ๋ฐ์ํจ ํ, ๋ฏผ๋ณ์ ์, ์ํํํ์ ์ ๋์ , ๊น๋ํฌ, ์ด๋์ "
"์ธ ์ ์จ, ์ํ์ฐ๊ตฌํ์ ๊น๊ท์ง ์จ, ์ผ๋ณธ์ธ์ธก ์ํ๊ฐ๋ก ๊ณ ๋ชฉ๋ฐฐ์ ์ธ ์์จ ๊ธฐํ ์ํ๊ฐ์ ๊ด๊ณ์๋ "
"์ธ์ฌ๋ฅผ ์ด์ฒญํ๊ณ ์์ผ ์ ๋ฌด์ด๊ฐ ์ดํ ํ๋ฌด๋น๊ตญ์๊ฐ ํํฉํ์ฌ ์ฐจ์ ๊ดํ ์์ํ๋ฅผ ๊ฐํ๋ฐ"
)
print(run("ํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:", hangul))
# ็พ่ก็พ่ก่ ่จๅ ็ธฝ็ฃๅบ ไบๆก์ผ๋ก ๆๅนด ๅๆฌก ้่จญ็ธฝ็ฃๅบ์์๋ ๆ้ฎฎ์ ๅจ ํ ็พ่ก์ ็ผ้์ ๅๅ ฑํ
# ็ฎ็์ผ๋ก ๆฑไบฌ์ ๅธๅ็พ่ก้ขๅฑ่ฆฝๆ๋ฅผ ่จชํ ์ฌ ๆฏๅนด ๆฅๆฌก ็พ่กๅฑ่ฆฝๆ๋ฅผ ้ํ ๆน้์ ๅ
งๅฎํ๊ณ ๆ
# ไบๅๅ
ญๆฅ ๅๅ ๅๆ ็ธฝ็ฃๅบ ็ฌฌไบๅ์ๅฎค์ ๆดๆฐธๅญ ๅพ, ้ไธ้ซ ๅญ, ๆธ็ตๅๆ์ ็ฒพๅคง้ก, ้ๆฆ็,
# ไผ้ๆฆฎ ๅค ่ซธๆฐ, ๆธ็ต็ก็ฉถๆ์ ้ๅฅ้ญ ๆฐ, ๆฅๆฌไบบๅด ๆธ็ตๅฎถ๋ก ้ซๆจๅนๆฐดๅค ๆฐ็ญ ๅ
ถไป ๆธ็ตๅฎถ์
# ้ไฟ์๋ ไบบไบ๋ฅผ ๆ่ซํ๊ณ ๅไน ๆฟๅ็ธฝ็ฃ ไปฅไธ ๅญธๅ็ถๅฑ่
๊ฐ ๆๅํ์ฌ ๆญค์ ้ํ ๅ่ญฐๆ๋ฅผ ้ํ๋ฐ
The outputs are shown as generated (each is one line). In the first, ๆๅ ด
came out as ํ์์ฅ where ํ์ฅ was meant. The second is the opening of the
1922 article; the original reads ็พ่กๅฑ่ฆฝ ่จๅ ็ธฝ็ฃๅบ ไบๆก์ผ๋ก ๆๅนด ๅๆฌก ้่จญ็ธฝ็ฃๅบ์์๋ ๆ้ฎฎ์ ๅจ ํ ็พ่ก์ ็ผ้์ ่ฃจ่ฃํ ๋ชฉ์ ์ผ๋ก ๆฑไบฌ์ ๅธๅ็พ่ก้ขๅฑ่ฆฝๆ๋ฅผ ๅฃํ ์ฌ ๆฏๅนด ไธๆฌก ..., so the output doubles a word at the
start (็พ่ก็พ่ก่) and writes same-sounding characters for several words and
names (ๅๅ ฑ, ่จช, ๆฅๆฌก, ็ฒพๅคง้ก).
Prompt format, as used by the SFT collator in the GitHub repo and by the chat template:
<|system|>{system}<|user|>{user}<|end_of_turn|>, with<|assistant|>{reply}<|end_of_turn|><|user|>{user}<|end_of_turn|>for each further turn. The system part is optional in the format, but it is where this checkpoint expects the task prompt. Each direction was trained with three wordings:ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:,๋ค์ ํ๋ฌธ์ ํ๊ธ๋ก ์ฎ๊ธฐ์์ค:,ํ์ ํ๊ธฐ๋ฅผ ํ๊ธ ๋ ์์ผ๋ก ๋ณํํ์์ค:for Hanja to Hangul, andํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:,๋ค์ ํ๊ธ ํ ์คํธ์ ์ ์ ํ ํ์๋ฅผ ์ถ๊ฐํ์์ค:,ํ์ ํ๊ธฐ๊ฐ ํ์ํ ๋ถ๋ถ์ ํ์๋ฅผ ๋ณ๊ธฐํ์์ค:for Hangul to Hanja.- The generation prompt is
<|assistant|>. A turn ends with<|end_of_turn|>(eos_token_id10). - Thinking. The chat template accepts
enable_thinking=True, which starts the reply at<|think|>, and the training data had reasoning examples. The model did not learn a separate reasoning step from them: started at<|think|>it answers as it does from<|assistant|>, or writes one sentence of rationale and ends the turn, and it never moves on to an<|assistant|>answer. Leave thinking off.
The wrapper splits the prompt at its last <|end_of_turn|>: everything up to
and including it is the encoder input, and the generation prompt after it starts
the decoder. A prompt without a generation prompt raises. forward() is not
wrapped and stays encoder-decoder.
Encoder-decoder generation
AutoModelForSeq2SeqLM loads the model itself. generate() is the standard
transformers one, with a KV cache; greedy decoding, beam search, and sampling
all work. The encoder takes the prompt up to <|end_of_turn|>, and the decoder
starts from <|assistant|> (decoder_start_token_id 9).
from transformers import AutoModelForSeq2SeqLM
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
prompt = (
"<|system|>ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:"
"<|user|>ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.<|end_of_turn|>"
)
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Tokenizer
tokenizer.json is a tokenizers-library build of the SentencePiece model
(spiece.model, kept here as the source), made by
scripts/build_hf_tokenizer.py in the GitHub repo. Calling the tokenizer does
not add BOS or EOS tokens, and special tokens written into the text are mapped
to their ids. It encodes like the SentencePiece wrapper the training code uses,
with one known difference: where two segmentations of a span have exactly the
same score (runs of digits, mostly), the same pieces can come out in a different
order.
Files
| File | Contents |
|---|---|
model.safetensors |
fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables |
config.json, generation_config.json |
model and generation config, with auto_map entries for the Auto classes |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
fast tokenizer (38400 pieces) and chat template |
spiece.model |
the SentencePiece model tokenizer.json was built from |
modeling_han2han.py, han2han_config.py |
modeling code from the GitHub repo at commit af1330e |
modeling_han2han.py (modeling_han2han_pytorch.py there) differs in two
places: the config import is relative, and the module-level register_han2han
import is replaced by the auto_map entries. han2han_config.py is the GitHub
copy at commit 7304c38; it has not changed since. config.json names
AutoModelForCausalLM under architectures so that transformers serve, which
looks that name up in the transformers namespace, can load the model.
Instruction tuning
- Pre-training: 35B tokens; the uniform average of its last five
checkpoints is
cadazar/han2han-pt. - Instruction tuning (
emnlp2026): 1.5B tokens from that average on instruction following, chain-of-thought reasoning, Hanja-Hangul article transcription, and summarization data, weights in bf16 (configs/it-muon-stage_1.yamlin the GitHub repo). - Continuation (
main): 30M more tokens on the same four task families in fp32, learning rate 3e-5, about 40k tokens per update, 762 updates.
How far the embedding row of <|end_of_turn|> sits from its pre-trained value
(L2 distance): 0.014 after step 2, 0.30 after step 3. Ordinary token rows moved
0.01 to 0.04 in step 3.
Citation
@inproceedings{han2han2026,
title = {Han2Han: Efficient Language-Specific Character Representation
through Script-Aware Pre-Training for Historical Text Analysis},
author = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}
License
Apache License 2.0, the same as the GitHub repo.
- Downloads last month
- 788