--- language: - bal license: apache-2.0 base_model: shahbakhsh/BalBERT tags: - token-classification - part-of-speech - pos-tagging - balochi - low-resource - universal-dependencies - transformers - pytorch - shahbakhsh datasets: - custom-balochi-ud metrics: - accuracy - f1 - matthews_correlation model-index: - name: BalPOS-v2 results: - task: type: token-classification name: Part-of-Speech Tagging dataset: name: Custom Balochi UD (CoNLL-U) type: custom-balochi-ud metrics: - type: accuracy name: Accuracy value: 0.8726 - type: f1 name: Macro F1 value: 0.7899 - type: f1 name: Weighted F1 value: 0.8722 - type: matthews_correlation name: Matthews Correlation Coefficient (MCC) value: 0.8526 - type: accuracy name: Balanced Accuracy value: 0.7881 - type: precision name: Macro Precision value: 0.7956 - type: recall name: Macro Recall value: 0.7881 widget: - text: "وتی فلسفہ" example_title: "Balochi Example 1" - text: "مئے ماتی زبان بلۆچی اِنت" example_title: "Balochi Example 2" - text: "آییءِ کتاب ماں شھرءَ انت" example_title: "Balochi Example 3" ---
# 🏷️ BalPOS v2 — Balochi Universal Part-of-Speech Tagger ### *State-of-the-Art Universal Dependencies POS Tagging for the Balochi Language*
[![Hugging Face Model](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-shahbakhsh%2FBalPOS-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black)](https://huggingface.co/shahbakhsh/BalPOS) [![GitHub Repository](https://img.shields.io/badge/GitHub-shah--bakhsh%2FBalPOS-181717?style=for-the-badge&logo=github)](https://github.com/shah-bakhsh/BalPOS) [![License: Apache 2.0](https://img.shields.io/badge/License-Apache_2.0-2ea44f?style=for-the-badge)](https://opensource.org/licenses/Apache-2.0) [![Base Model](https://img.shields.io/badge/Base%20Model-shahbakhsh%2FBalBERT-blue?style=for-the-badge)](https://huggingface.co/shahbakhsh/BalBERT)
| **Accuracy** | **Weighted F1** | **MCC** | **Macro F1** | |:---:|:---:|:---:|:---:| | **87.26 %** | **87.22 %** | **0.8526** | **78.99 %** |
--- ## 💎 Model Highlights & Architecture **BalPOS v2** is a transformer-based token-classification model fine-tuned from **[`shahbakhsh/BalBERT`](https://huggingface.co/shahbakhsh/BalBERT)** specifically for **Universal Dependencies (UD) Part-of-Speech tagging** on the Balochi language (`bal`). Developed by **[Shah Bakhsh](https://huggingface.co/shahbakhsh)**, BalPOS v2 provides the foundational syntactic layer for the Balochi NLP research ecosystem. ### Key Architectural Attributes - **Base Architecture:** Pretrained Transformer (`shahbakhsh/BalBERT`) - **Task:** Token Classification (Universal POS Tagging) - **Tagset:** 16 Universal Dependencies (UPOS) tags - **Target Language:** Balochi (`bal`) - **Validation Protocol:** 5-Fold Cross-Validation over 85% data + 15% Untouched Final Hold-Out Test Set - **Deterministic Training:** Fixed random seed (`42`), cuDNN deterministic execution enabled --- ## 🎯 Benchmark Performance (Final Hold-Out Test) The evaluation results below reflect performance on the **15% final hold-out test set**, which was never seen during hyperparameter optimization or fold selection:
| Metric | Exact Score | Percentage | Status | | :--- | :---: | :---: | :---: | | **Accuracy** | `0.8726` | **87.26%** | 🥇 Primary Benchmark | | **Weighted F1** | `0.8722` | **87.22%** | ⚡ Distribution Weighted | | **Weighted Precision** | `0.8741` | **87.41%** | ⚡ Distribution Weighted | | **Weighted Recall** | `0.8726` | **87.26%** | ⚡ Distribution Weighted | | **Matthews Correlation Coefficient (MCC)** | `0.8526` | **0.8526** | 📐 Multi-Class Quality | | **Macro F1** | `0.7899` | **78.99%** | ⚖️ Unweighted Class Mean | | **Macro Precision** | `0.7956` | **79.56%** | ⚖️ Unweighted Class Mean | | **Macro Recall / Balanced Accuracy** | `0.7881` | **78.81%** | ⚖️ Unweighted Class Mean |
--- ## ⚙️ Hyperparameters & Model Selection Model selection was performed via **5-fold cross-validation** over the development split following randomized search hyperparameter optimization: ```json { "learning_rate": 1e-05, "batch_size": 8, "epochs": 15, "weight_decay": 0.01, "warmup_ratio": 0.05, "gradient_accumulation_steps": 1, "seed": 42, "best_fold": "Fold 5" } ``` --- ## 🚀 Usage & Quickstart ### Option 1: High-Level `pipeline` ```python from transformers import pipeline # Load BalPOS v2 pipeline from Hugging Face tagger = pipeline( task="token-classification", model="shahbakhsh/BalPOS", aggregation_strategy="simple" ) # Tag Balochi sentence text = "وتی فلسفہ" results = tagger(text) for entity in results: print(f"Token: {entity['word']:<15} | Tag: {entity['entity_group']:<8} | Score: {entity['score']:.4f}") ``` ### Option 2: Direct PyTorch / Transformers API ```python import torch from transformers import AutoTokenizer, AutoModelForTokenClassification model_id = "shahbakhsh/BalPOS" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForTokenClassification.from_pretrained(model_id) model.eval() sentence = "وتی فلسفہ" inputs = tokenizer(sentence, return_tensors="pt") with torch.no_grad(): logits = model(**inputs).logits predictions = torch.argmax(logits, dim=-1)[0] tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0]) for token, pred_id in zip(tokens, predictions): if token not in [tokenizer.cls_token, tokenizer.sep_token, tokenizer.pad_token]: tag = model.config.id2label[pred_id.item()] print(f"{token:<15} -> {tag}") ``` --- ## 📊 Universal POS Tagset (16 Labels) BalPOS v2 maps Balochi tokens to all 16 UPOS categories: | Tag | Name | Description | Example Tokens | | :---: | :--- | :--- | :--- | | **ADJ** | Adjective | Modifies a noun | گران, شَر | | **ADP** | Adposition | Preposition / Postposition | گۆں, ماں | | **ADV** | Adverb | Modifies verb/adj | انّۆ, سک | | **AUX** | Auxiliary | Auxiliary or copular verb | اِنت, بوت | | **CCONJ** | Coordinating Conjunction | Connects words or clauses | ءُ, یا | | **DET** | Determiner | Demonstrative or article | اے, آ | | **INTJ** | Interjection | Exclamation | واہ, ھۆ | | **NOUN** | Noun | Common noun | مارچ, کتاب | | **NUM** | Numeral | Cardinal/ordinal number | یک, دۆ | | **PART** | Particle | Function word | مئے | | **PRON** | Pronoun | Personal/possessive pronoun | من, تئو, وتی | | **PROPN** | Proper Noun | Name of entity | بلوچستان, کراچی | | **PUNCT** | Punctuation | Punctuation marks | . , ؟ ! | | **SCONJ** | Subordinating Conjunction | Subordinating connector | کہ, پرچا کہ | | **VERB** | Verb | Action/state verb | رَوت, گوشیت | | **X** | Other | Foreign or unclassified token | — | --- ## 📂 Corpus & Dataset Information - **Corpus:** Custom Balochi UD (CoNLL-U) Corpus - **Total Corpus Size:** 774 sentences / 14,852 tokens - **Development Partition (85%):** Used for 5-Fold Cross-Validation & HPO - **Test Partition (15%):** Untouched hold-out test set for final reporting --- ## ⚠️ Limitations & Ethical Considerations 1. **Single-Source Corpus:** Annotations and vocabulary are inherited from the source CoNLL-U corpus. 2. **Rare Tags Variance:** Infrequent tags (such as `INTJ` or `X`) have higher variance in per-class evaluation. 3. **Dialectal Scope:** Evaluated primarily on standard orthographic conventions present in the training set. --- ## 🗺️ Balochi NLP Ecosystem Roadmap BalPOS v2 is part of an ongoing open-source initiative for Balochi NLP led by [Shah Bakhsh](https://huggingface.co/shahbakhsh): - 🟢 **[`shahbakhsh/BalBERT`](https://huggingface.co/shahbakhsh/BalBERT)** — Pretrained Masked Language Model *(Released)* - 🟢 **[`shahbakhsh/BalPOS`](https://huggingface.co/shahbakhsh/BalPOS)** — Universal POS Tagger *(Released)* - 🟡 **BalNER** — Named Entity Recognition for Balochi *(In Development)* - 🟡 **BalMorph** — Morphological Feature Tagging *(In Development)* - 🟡 **BalParser** — Universal Dependency Parser *(In Development)* --- ## 📜 Citation & Attribution If you use **BalPOS v2** or **BalBERT**, please cite this repository: ```bibtex @misc{balpos_v2_2026, title = {{BalPOS v2: Balochi Universal Part-of-Speech Tagger}}, author = {Shah Bakhsh}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/shahbakhsh/BalPOS}}, note = {Fine-tuned from shahbakhsh/BalBERT on Balochi UD Corpus} } ``` --- ## 📄 License Distributed under the **[Apache License 2.0](LICENSE)**.