| --- |
| language: |
| - si |
| license: mit |
| tags: |
| - sinhala |
| - nlp |
| - spellcheck |
| - typo-detection |
| - seq2seq |
| - sequence-labeling |
| - pytorch |
| metrics: |
| - accuracy |
| --- |
| |
|  |
|
|
| # sinlib: Pretrained Sinhala NLP and Spellchecking Models |
|
|
| This repository hosts the official pretrained model checkpoints, vocabularies, and statistical datasets for the **`sinlib`** library—a comprehensive Sinhala NLP toolkit. |
|
|
| ## Models & Artifacts Hosted |
|
|
| This repository contains the following files loaded dynamically by `sinlib.spellcheck.TypoDetector`: |
|
|
| * **`bigru_detector.pt`**: A bidirectional GRU sequence labeling model (`BiGRUSequenceLabeler`) trained to detect spelling errors and character substitutions at the akshara level. |
| * **`bigru_corrector.pt`**: A sequence-to-sequence bidirectional GRU encoder-decoder model with Attention (`BiGRUSeq2Seq`) that performs generative character/akshara corrections. |
| * **`akshara_vocab.json`**: Vocabulary mappings mapping Sinhala phonological units (aksharas) and basic punctuation to token IDs for the neural models. |
| * **`akshara_ngram.json`**: Stored counts and vocabularies used by the statistical Trigram model (`AksharaNGram`) for likelihood evaluation. |
| * **`news_unigrams.json` & `news_bigrams.json`**: Pre-calculated unigram and bigram word-level frequencies extracted from Sinhala news corpora, used for context-aware candidate re-ranking (Stupid Backoff). |
| * **`dictionary.npy`**: The canonical baseline dictionary containing ~61K valid Sinhala words. |
| * **`ngram_probs.npy`**: N-gram probabilities used by the default `PreTrainedTokenizer` fallback checks. |
| |
| Additional models, configuration files, and vocabulary resources required by the `sinlib` package are also included in this repository. |
| |
| --- |
| |
| ## Intended Use |
| |
| These weights are designed to be loaded directly through the `sinlib` Python package. |
| |
| ### Installation |
| ```bash |
| pip install sinlib |
| ``` |
| |
| ### Example Inference (Spellcheck) |
| ```python |
| from sinlib.spellcheck import TypoDetector |
| |
| # Automatically downloads and caches the model files from this HF repository |
| detector = TypoDetector.from_pretrained("Ransaka/sinlib") |
| |
| # Run spellcheck (with punctuation preservation & unicode normalization) |
| sentence = "කොළඹ වරායේ සිට බස්නහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරන්." |
| corrected = detector(sentence) |
| |
| print(corrected) |
| # Output: "කොළඹ වරායේ සිට බස්නාහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරින්." |
| ``` |
| |
| ### Typo Detection Fallback |
| If PyTorch is not installed in the target environment, `sinlib` falls back automatically to the statistical N-Gram model (`akshara_ngram.json`) to perform spelling checks. |
| ```python |
| # Check word suspicion level |
| is_typo = detector.is_word_suspicious("පසලට") # True |
| ``` |
| |
| For more details, usage examples, and API references, please refer to the official documentation: |
| |
| https://sinlib.readthedocs.io/en/latest/ |