--- license: mit language: - bal tags: - dependency-parsing - pos-tagging - biaffine - pytorch - balochi pipeline_tag: text-classification model_type: biaffine-dependency-parser --- # Balochi Dependency Parser 🇧🇦 A from-scratch **biaffine dependency parser** for the Balochi language — an Iranian language spoken by ~8 million people. Word-level and character-level representations are learned jointly from a 27,000+ sentence treebank — **no pretrained embeddings, no transfer learning** — with joint **POS tagging + dependency parsing**, **MST decoding**, and an **honest, calibrated confidence** system. This repo hosts the model weights in **safetensors** format (no pickle — safe to load anywhere). The companion `config.json` + `vocab.json` let you rebuild the exact architecture, and `calibration.json` provides the Platt-scaled confidence coefficients. ## Model details | Property | Value | |----------|-------| | Architecture | Biaffine parser (Dozat & Manning 2017), from scratch | | Encoder | 4-layer BiLSTM (hidden 300) over word (150-d) + char (64-d) embeddings | | Vocab | 160,875 word types / 268 chars | | Decoding | Chu-Liu-Edmonds MST with UD single-root constraint | | Extras | 12-rule UD constraint layer, rule-based morphology, enhanced DEPS | | Parameters | 33,811,073 (56 tensors) | ### Accuracy (official test split, 2,788 sentences) | Metric | Value | |--------|-------| | UAS | **94.31%** | | LAS | **91.98%** (with UD constraints) | | POS accuracy | **98.20%** | Confidence calibration: ECE 0.135 → 0.011 on test (Platt scaling). At a calibrated gate of ≥ 0.95, 99.74% of emitted parses are exact. > The treebank and model use **Perso-Arabic script** — Latin-script Balochi > is not supported. ## Usage This is a **custom architecture**, not a `transformers` model — there is no `AutoModel`. Load it with the project's own loader (`balochi_parser.checkpoint.load_checkpoint_safetensors`), which reads `model.safetensors` + `config.json` + `vocab.json` from this repo. ### 1. Clone the source repo and install ```bash git clone balochi-parser cd balochi-parser/backend pip install -e ".[dev,serving]" pip install safetensors ``` ### 2. Download the weights from this repo ```bash cd balochi-parser/backend # put the 4 files into a folder, e.g. hf_export/ # (or use huggingface_hub: huggingface-cli download /balochi-parser --local-dir new_model) ``` ### 3. Load and parse ```python import torch from balochi_parser.checkpoint import load_checkpoint_safetensors from balochi_parser.predict import parse_sentence from balochi_parser.tokenizer import pretokenize model, vocab, config, _ = load_checkpoint_safetensors("new_model") model.eval() tokens = pretokenize("ائی مرد بازارءَ شت", vocab["word2idx"]) rows = parse_sentence(model, vocab, tokens, torch.device("cpu")) for row in rows: print(f"{row['id']}\t{row['form']}\t{row['upos']}\t{row['head']}\t{row['deprel']}") ``` Output (CoNLL-U-style): ``` 1 ائی PRON 5 nsubj 2 مرد NOUN 5 obj 3 بازار NOUN 5 obl 4 ءَ ADP 3 case 5 شت VERB 0 root ``` Each token row also carries `lemma`, `feats`, `deps`, `misc`, and a per-token confidence `conf` (P(arc) × P(label)). `calibration.json` sits next to the weights, so confidence calibration is applied automatically. ### CLI (from the source repo) ```bash cd backend python -m balochi_parser.predict --ckpt new_model/model.safetensors \ --text "ائی مرد بازارءَ شت" --show_conf ``` ### Expand the vocabulary without retraining Dialect words absent from training get their own embedding at load time (vocab grows, existing embeddings stay bit-for-bit). Works on the safetensors export too: ```python model, vocab, config, _ = load_checkpoint_safetensors( "new_model", extra_vocab="data/external_text/vocab_makrani.txt" ) # or through the shared loader (accepts the file or the directory): # load_checkpoint("new_model/model.safetensors", extra_vocab="my_words.txt") ``` ```bash # CLI — the --ckpt flag accepts the .safetensors file directly: python -m balochi_parser.predict --ckpt new_model/model.safetensors \ --extra_vocab my_words.txt --text "تو کُجئیگ ئے" --show_conf ``` ## Files | File | Description | |------|-------------| | `model.safetensors` | Model weights (56 tensors, 135 MB) | | `config.json` | Architecture config + vocab sizes | | `vocab.json` | `word2idx` / `char2idx` / `upos2idx` / `deprel2idx` | | `calibration.json` | Platt-scaling coefficients (confidence calibration) | | `README.md` | This file | ## Why safetensors? The original checkpoint (`new_model/best_model.pt`) is a `torch.save` pickle. `model.safetensors` is the same weights in a pickle-free, size-safe format — the exact export is verified bit-for-bit identical (`scripts/export_safetensors.py --verify`). ## Training data & citation - **Treebank**: 27,824 gold sentences (22,250 train / 2,786 dev / 2,788 test), augmented to 91,170 for training (gold + pseudo-labeled + BNER). - **Sources**: Balochi Academy texts (novels, folktales, proverbs, poetry, articles) + public Southern/Makrani sources, incl. 17,854 validated Makrani sentences and a 22,559-word Makrani wordlist. If you use this model in research, please cite: ```bibtex @software{balochi_parser2026, title={Balochi Dependency Parser}, year={2026}, description={A from-scratch biaffine dependency parser for Balochi}, } ``` ## License MIT — see the source repo's `LICENSE`. --- Built with ❤️ for the Balochi language community.