| --- |
| language: multilingual |
| license: apache-2.0 |
| tags: |
| - language-identification |
| - text-classification |
| - programming-language-detection |
| - spam-detection |
| - toxicity-detection |
| - byte-level |
| - multi-task |
| - lightweight |
| pipeline_tag: text-classification |
| --- |
| |
| # Glyph — Multi-Task Byte-Level Text Classifier |
|
|
| **4M parameters** · **5 tasks** · **No tokenizer needed** · **Runs on CPU** |
|
|
| Glyph is an ultra-compact multi-task text classification model that operates directly on raw UTF-8 bytes. A single shared backbone serves 5 classification heads simultaneously. |
|
|
| ## Tasks & Performance |
|
|
| | Task | Labels | Val Accuracy | Description | |
| |------|--------|-------------|-------------| |
| | `lang_id` | 351 | 94.3% | Language identification | |
| | `prog_lang` | 30 | 94.4% | Programming language detection | |
| | `spam` | 2 | 96.1% | Spam vs ham classification | |
| | `toxic` | 2 | 67.3% | Toxicity detection (multilingual) | |
|
|
| **Average validation accuracy: 88.0%** |
|
|
| ## Architecture |
|
|
| Glyph uses a byte-level hybrid architecture inspired by [CommonLingua](https://huggingface.co/PleIAs/CommonLingua): |
|
|
| - **Input**: Raw UTF-8 bytes (no tokenizer), padded to 512 bytes |
| - **Trigram hash embedding**: Polynomial rolling hash of byte 3-grams → 8192-bucket embedding table |
| - **Byte unigram embedding**: Standard embedding for individual bytes |
| - **4× Conv1D blocks**: Causal convolutions + BatchNorm + GELU + residual |
| - **2× Bidirectional attention**: Multi-head self-attention with RoPE |
| - **Global average pooling** → per-task classification heads |
|
|
| Total: ~4M shared parameters + small per-task linear heads. |
|
|
| ## Usage |
|
|
| ```python |
| import torch, importlib, sys |
| from huggingface_hub import hf_hub_download |
| |
| # Download model + weights |
| model_py_path = hf_hub_download("ThingAI/Glyph", "model.py") |
| ckpt_path = hf_hub_download("ThingAI/Glyph", "model.pt") |
| |
| # Load model definition |
| import importlib.util |
| spec = importlib.util.spec_from_file_location("model", model_py_path) |
| mod = importlib.util.module_from_spec(spec) |
| spec.loader.exec_module(mod) |
| |
| # Load weights |
| ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False) |
| model = mod.MultiTaskLID(ckpt["task_configs"]).eval() |
| model.load_state_dict(ckpt["model"]) |
| |
| # Predict |
| def predict(text, task="lang_id", top_k=3): |
| raw = text.encode("utf-8")[:512] |
| byte_ids = list(raw) + [0] * (512 - len(raw)) |
| inp = torch.tensor([byte_ids], dtype=torch.long) |
| with torch.no_grad(): |
| logits = model(inp, task)["logits"][0] |
| probs = torch.softmax(logits, dim=-1) |
| idx2label = {v: k for k, v in ckpt["label_maps"][task].items()} |
| k = min(top_k, len(probs)) |
| topk = probs.topk(k) |
| return [(idx2label[topk.indices[i].item()], topk.values[i].item()) for i in range(k)] |
| |
| # Examples |
| print(predict("La pizza napoletana è patrimonio UNESCO", task="lang_id")) |
| print(predict("def foo(x): return x + 1", task="prog_lang")) |
| print(predict("You won a FREE iPhone!!!", task="spam")) |
| ``` |
|
|
| ## Quick predict script |
|
|
| ```bash |
| python predict.py --text "Ciao, come stai?" |
| # lang_id: ita (98.2%) |
| |
| python predict.py --task prog_lang --text "fn main() { println!("hello"); }" |
| # prog_lang: Rust (97.1%) |
| |
| python predict.py --task spam --text "URGENT: Click here to win!" |
| # spam: spam (99.5%) |
| ``` |
|
|
| ## Training Data |
|
|
| | Task | Dataset | Samples | |
| |------|---------|---------| |
| | Language ID | [PleIAs/CommonLingua-Train](https://huggingface.co/datasets/PleIAs/CommonLingua-Train) | 200K (capped) | |
| | Programming Language | [cakiki/rosetta-code](https://huggingface.co/datasets/cakiki/rosetta-code) | ~26K | |
| | Spam Detection | [ucirvine/sms_spam](https://huggingface.co/datasets/ucirvine/sms_spam) | 5.5K | |
| | Toxicity | [textdetox/multilingual_toxicity_dataset](https://huggingface.co/datasets/textdetox/multilingual_toxicity_dataset) | ~45K | |
|
|
| ## Key Design Decisions |
|
|
| - **Byte-level input**: No tokenizer means it works on any language, script, or encoding without preprocessing |
| - **Trigram hashing**: +1.2 F1 over unigram-only baseline, acts as regularization via hash collisions |
| - **Multi-task learning**: Shared backbone learns universal text representations; per-task heads are tiny (~130K params each) |
| - **No attention masking**: Bidirectional attention for classification (not causal) |
| - **OneCycleLR**: Fast convergence in few epochs |
|
|
| ## Limitations |
|
|
| - Designed for paragraph-level classification (50+ bytes). Short texts (<20 bytes) may be unreliable |
| - Toxicity detection accuracy is lower than specialized models (trained on limited multilingual data) |
| - Programming language detection trained on Rosetta Code samples which have a specific style |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{glyph2026, |
| author = {ThingAI}, |
| title = {Glyph: Multi-Task Byte-Level Text Classifier}, |
| year = {2026}, |
| url = {https://huggingface.co/ThingAI/Glyph} |
| } |
| ``` |
|
|
| ## License |
|
|
| Apache 2.0 |
|
|