--- language: - si license: mit tags: - sinhala - nlp - spellcheck - typo-detection - seq2seq - sequence-labeling - pytorch metrics: - accuracy --- ![image](https://cdn-uploads.huggingface.co/production/uploads/60f2e10dadf471cbdf8bb661/MHw89go_R_-fcceMu2zRT.png) # sinlib: Pretrained Sinhala NLP and Spellchecking Models This repository hosts the official pretrained model checkpoints, vocabularies, and statistical datasets for the **`sinlib`** library—a comprehensive Sinhala NLP toolkit. ## Models & Artifacts Hosted This repository contains the following files loaded dynamically by `sinlib.spellcheck.TypoDetector`: * **`bigru_detector.pt`**: A bidirectional GRU sequence labeling model (`BiGRUSequenceLabeler`) trained to detect spelling errors and character substitutions at the akshara level. * **`bigru_corrector.pt`**: A sequence-to-sequence bidirectional GRU encoder-decoder model with Attention (`BiGRUSeq2Seq`) that performs generative character/akshara corrections. * **`akshara_vocab.json`**: Vocabulary mappings mapping Sinhala phonological units (aksharas) and basic punctuation to token IDs for the neural models. * **`akshara_ngram.json`**: Stored counts and vocabularies used by the statistical Trigram model (`AksharaNGram`) for likelihood evaluation. * **`news_unigrams.json` & `news_bigrams.json`**: Pre-calculated unigram and bigram word-level frequencies extracted from Sinhala news corpora, used for context-aware candidate re-ranking (Stupid Backoff). * **`dictionary.npy`**: The canonical baseline dictionary containing ~61K valid Sinhala words. * **`ngram_probs.npy`**: N-gram probabilities used by the default `PreTrainedTokenizer` fallback checks. Additional models, configuration files, and vocabulary resources required by the `sinlib` package are also included in this repository. --- ## Intended Use These weights are designed to be loaded directly through the `sinlib` Python package. ### Installation ```bash pip install sinlib ``` ### Example Inference (Spellcheck) ```python from sinlib.spellcheck import TypoDetector # Automatically downloads and caches the model files from this HF repository detector = TypoDetector.from_pretrained("Ransaka/sinlib") # Run spellcheck (with punctuation preservation & unicode normalization) sentence = "කොළඹ වරායේ සිට බස්නහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරන්." corrected = detector(sentence) print(corrected) # Output: "කොළඹ වරායේ සිට බස්නාහිර දෙසින් නාවික සැතපුම් 17ක් පමණ දුරින්." ``` ### Typo Detection Fallback If PyTorch is not installed in the target environment, `sinlib` falls back automatically to the statistical N-Gram model (`akshara_ngram.json`) to perform spelling checks. ```python # Check word suspicion level is_typo = detector.is_word_suspicious("පසලට") # True ``` For more details, usage examples, and API references, please refer to the official documentation: https://sinlib.readthedocs.io/en/latest/