Han2Han PT

The pre-trained (PT) checkpoint of Han2Han, a 169M-parameter encoder-decoder model that learns script-invariant representations of Korean text: a document written in Hanja and its Hangul transcription land at the same point in embedding space. The recipe (jamo and character-level embedding fusion, morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described in the paper, accepted to Findings of EMNLP 2026.

This is the starting point for fine-tuning, and the first of three checkpoints:

Repo Stage
cadazar/han2han-pt (this one) pre-training, 35B tokens
cadazar/han2han-it instruction tuning from these weights
cadazar/han2han-rl reinforcement learning on transcription in both directions (Hangul to Hanja and Hanja to Hangul), from han2han-it

These weights are an average

The weights are the uniform average of the last five checkpoints of the pre-training run (steps 6347, 6441, 6536, 6630 and 6725, saved every 0.5B tokens from 33B to 35B), not the final checkpoint alone. The average is what the instruction tuning started from, chosen after comparing starting points, so it is the pre-trained model that han2han-it and han2han-rl descend from. The learning rate had nearly finished its cooldown over those steps: no weight differs from the final checkpoint by more than 0.0015.

Intended use

Fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike.

It is not a chat model. The tokenizer and the chat template are the same as in the other two repos, but pre-training never used the chat tokens (<|system|>, <|user|>, <|assistant|>, <|think|>, <|end_of_turn|>): their embedding rows are still nearly one shared vector here, as for every other token pre-training never saw. For generation from a prompt use han2han-it or han2han-rl.

Usage

Runtime requirements: torch and transformers (tested with torch 2.14.1 and transformers 5.18.0 on CPU). The model class ships in this repo, so loading needs trust_remote_code=True; the tokenizer itself runs no custom code, but without the flag AutoTokenizer stops to ask.

output_sentence_embeddings=True returns the mean-pooled encoder states, one 640-dimensional vector per input.

import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

repo = "cadazar/han2han-pt"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()

texts = [
    "ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ํšŒ์žฅ์„ ์ผ์ˆœํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ์ผ๋ณธ์ธ์ธก ํ™”๊ฐ€๊นŒ์ง€๋„ ์ผ์ธ๋„ ๋ฐœ๊ฒฌํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ๅ—็•ต๋‚˜ ๅ››ๅ›ๅญ์—์„œ๋Š” ็ ดๅขจ์˜ ๅฆ™ๆณ•์œผ๋กœ ็™ฝ้›ช์„ ่ฑกๅพตํ•  ์ˆ˜ ์žˆ๋‹ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
    embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9315, 0.8253],
#         [0.9315, 1.0000, 0.7910],
#         [0.8253, 0.7910, 1.0000]])

The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.

Tokenizer

tokenizer.json is a tokenizers-library build of the SentencePiece model (spiece.model, kept here as the source), made by scripts/build_hf_tokenizer.py in the GitHub repo. Calling the tokenizer does not add BOS or EOS tokens, and special tokens written into the text are mapped to their ids. It encodes like the SentencePiece wrapper the training code uses, with one known difference: where two segmentations of a span have exactly the same score (runs of digits, mostly), the same pieces can come out in a different order.

Files

File Contents
model.safetensors fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables
config.json, generation_config.json model and generation config, with auto_map entries for the Auto classes
tokenizer.json, tokenizer_config.json, chat_template.jinja fast tokenizer (38400 pieces) and chat template, the same as in han2han-it
spiece.model the SentencePiece model tokenizer.json was built from
modeling_han2han.py, han2han_config.py modeling code from the GitHub repo at commit af1330e, the same files as in han2han-it

Citation

@inproceedings{han2han2026,
  title     = {Han2Han: Efficient Language-Specific Character Representation
               through Script-Aware Pre-Training for Historical Text Analysis},
  author    = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}

License

Apache License 2.0, the same as the GitHub repo.

Downloads last month
31
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support