Download README.md from ChrisMoe/template: direct link, hf CLI and curl.
- Browser
- Download file 4.84 kB
-
https://huggingface.co/ChrisMoe/template/resolve/main/README.md
- Command line
-
hf download hf://ChrisMoe/template/README.md
-
curl -L -o README.md https://huggingface.co/ChrisMoe/template/resolve/main/README.md
language: zh
library_name: onnx
tags:
- handwriting
- chinese
- hanzi
- stroke-similarity
- education
- onnx
HanziFlow stroke encoder
A small neural network (about 69,000 parameters) that scores how well one drawn Chinese stroke matches one template stroke. It is part of HanziFlow, a Mandarin learning app, and sits inside a stroke-matching pipeline that tells a learner which stroke or part of a character is missing, extra, backwards or out of order.
It is not a character recognizer: it compares strokes against a known target character.
Files
| File | What it is |
|---|---|
stroke_encoder.onnx |
The trained network for serving with onnxruntime (CPU is enough) |
stroke_encoder.pt |
The same network as PyTorch weights plus settings. Load with torch.load(path, weights_only=True) |
template_embeddings.npz |
Precomputed 64-number embeddings of every template stroke of the 256 supported characters (emb_<i>; chars gives the order) |
calibration.json |
scale (embedding distance to matcher cost), tau (0.6, match threshold), good (0.35, neat match), alpha_recommended (0.5, blend with the geometric cost) |
learned_cost.py |
Small helper that loads the files above and returns a cost function for the matcher |
How it works
- Input: one stroke resampled to 32 points in the Make Me a Hanzi frame (1024-unit box, y up), turned into a
(32, 6)feature array (position, direction of travel, position relative to the stroke start). SeeLearnedCost.features. - Output: a 64-number embedding. The cost is the distance between a drawn stroke's embedding and a template stroke's embedding, divided by
scale. Below about 0.35 is a neat match; a stroke is paired with a template stroke only below 0.6. - Recommended use: blend 50/50 with a geometric cost (
alpha = 0.5). The matcher itself (pairing, order and direction checks) is separate code and is not included.
from learned_cost import LearnedCost
lc = LearnedCost() # reads the data files in this folder
costs = lc.costs("好", strokes) # strokes: (n, 32, 2) array in the 1024-unit frame; costs[i, j] = drawn stroke i vs template stroke j
Training data
- Template strokes: Make Me a Hanzi stroke centre-lines for 256 characters.
- Simulated handwriting: a simulator distorts the template strokes (shrinks, shifts, rotates, wobbles, trims ends) at three sloppiness levels; about 30 simulated writings per training character; 179 training, 38 validation and 39 test characters (test characters are never seen in training); 20 epochs.
- Real handwriting: none. This model was trained on simulated handwriting only.
Evaluation
Measured on the 39 unseen test characters with fresh simulated writings (8 per character and level). Same writings for every column.
| Test | geometric (notebook 1) | learned only | blend 50/50 |
|---|---|---|---|
| right template stroke chosen (medium) | 99.4% | 99.7% | 99.8% |
| right template stroke chosen (heavy) | 94.4% | 95.9% | 96.7% |
| correct writing accepted (medium) | 99.7% | 100.0% | 99.7% |
| correct writing accepted (heavy) | 91.7% | 92.9% | 93.9% |
| wrong character accepted (lower is better) | 0.9% | 0.0% | 0.4% |
| one dropped stroke diagnosed exactly | 97.4% | 99.0% | 99.3% |
These numbers come from simulated handwriting made by the same simulator that produced the training data. Small differences between columns are within noise. They show how the three variants compare on simulated data; they do not show how the model performs on real learners.
Limitations
- Supports only the 256 characters in
template_embeddings.npz; a new character needs its template embeddings computed. - Not evaluated on children's writing or on any specific group of learners. Not evaluated on real handwriting.
- Judges stroke shape and position only; stroke order and direction are checked separately by the matcher.
- Near-identical characters (for example 牛 and 午) cannot be told apart by this model alone.
- A learning aid that gives feedback, not a tool for grading or any high-stakes decision.
License and attribution
- Template strokes come from Make Me a Hanzi:
dictionary.txtis LGPL-3.0-or-later andgraphics.txtis under the Arphic Public License.template_embeddings.npzcontains numbers derived from those glyph shapes; check the Arphic Public License terms before choosing a license for derived data. - Choose a license for the code and weights and add it to the
license:line at the top of this file.
Citation
HanziFlow (Smart Chinese Learning Hub), School of Information Technology, King Mongkut's University of Technology Thonburi, 2026.