Glyph — Multi-Task Byte-Level Text Classifier
4M parameters · 5 tasks · No tokenizer needed · Runs on CPU
Glyph is an ultra-compact multi-task text classification model that operates directly on raw UTF-8 bytes. A single shared backbone serves 5 classification heads simultaneously.
Tasks & Performance
| Task | Labels | Val Accuracy | Description |
|---|---|---|---|
lang_id |
351 | 94.3% | Language identification |
prog_lang |
30 | 94.4% | Programming language detection |
spam |
2 | 96.1% | Spam vs ham classification |
toxic |
2 | 67.3% | Toxicity detection (multilingual) |
Average validation accuracy: 88.0%
Architecture
Glyph uses a byte-level hybrid architecture inspired by CommonLingua:
- Input: Raw UTF-8 bytes (no tokenizer), padded to 512 bytes
- Trigram hash embedding: Polynomial rolling hash of byte 3-grams → 8192-bucket embedding table
- Byte unigram embedding: Standard embedding for individual bytes
- 4× Conv1D blocks: Causal convolutions + BatchNorm + GELU + residual
- 2× Bidirectional attention: Multi-head self-attention with RoPE
- Global average pooling → per-task classification heads
Total: ~4M shared parameters + small per-task linear heads.
Usage
import torch, importlib, sys
from huggingface_hub import hf_hub_download
# Download model + weights
model_py_path = hf_hub_download("ThingAI/Glyph", "model.py")
ckpt_path = hf_hub_download("ThingAI/Glyph", "model.pt")
# Load model definition
import importlib.util
spec = importlib.util.spec_from_file_location("model", model_py_path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
# Load weights
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model = mod.MultiTaskLID(ckpt["task_configs"]).eval()
model.load_state_dict(ckpt["model"])
# Predict
def predict(text, task="lang_id", top_k=3):
raw = text.encode("utf-8")[:512]
byte_ids = list(raw) + [0] * (512 - len(raw))
inp = torch.tensor([byte_ids], dtype=torch.long)
with torch.no_grad():
logits = model(inp, task)["logits"][0]
probs = torch.softmax(logits, dim=-1)
idx2label = {v: k for k, v in ckpt["label_maps"][task].items()}
k = min(top_k, len(probs))
topk = probs.topk(k)
return [(idx2label[topk.indices[i].item()], topk.values[i].item()) for i in range(k)]
# Examples
print(predict("La pizza napoletana è patrimonio UNESCO", task="lang_id"))
print(predict("def foo(x): return x + 1", task="prog_lang"))
print(predict("You won a FREE iPhone!!!", task="spam"))
Quick predict script
python predict.py --text "Ciao, come stai?"
# lang_id: ita (98.2%)
python predict.py --task prog_lang --text "fn main() { println!("hello"); }"
# prog_lang: Rust (97.1%)
python predict.py --task spam --text "URGENT: Click here to win!"
# spam: spam (99.5%)
Training Data
| Task | Dataset | Samples |
|---|---|---|
| Language ID | PleIAs/CommonLingua-Train | 200K (capped) |
| Programming Language | cakiki/rosetta-code | ~26K |
| Spam Detection | ucirvine/sms_spam | 5.5K |
| Toxicity | textdetox/multilingual_toxicity_dataset | ~45K |
Key Design Decisions
- Byte-level input: No tokenizer means it works on any language, script, or encoding without preprocessing
- Trigram hashing: +1.2 F1 over unigram-only baseline, acts as regularization via hash collisions
- Multi-task learning: Shared backbone learns universal text representations; per-task heads are tiny (~130K params each)
- No attention masking: Bidirectional attention for classification (not causal)
- OneCycleLR: Fast convergence in few epochs
Limitations
- Designed for paragraph-level classification (50+ bytes). Short texts (<20 bytes) may be unreliable
- Toxicity detection accuracy is lower than specialized models (trained on limited multilingual data)
- Programming language detection trained on Rosetta Code samples which have a specific style
Citation
@misc{glyph2026,
author = {ThingAI},
title = {Glyph: Multi-Task Byte-Level Text Classifier},
year = {2026},
url = {https://huggingface.co/ThingAI/Glyph}
}
License
Apache 2.0