Supernova Nepali Normalizer V2
Supernova Nepali Normalizer V2 is an independent, lightweight, rule-based text normalization system designed for Nepali and Nepali-English mixed text.
Features
- Unicode NFC normalization
- Zero-width character cleanup
- Whitespace normalization
- Newline normalization
- Tab normalization
- Punctuation normalization
- Nepali + English mixed-text preservation
- URL and email preservation
- Emoji preservation
- Very fast CPU processing
Validation
The normalizer was validated on 9,112 text samples containing approximately 936K characters.
Validation results:
- UTF-8 failures: 0
- NFC failures: 0
- Empty outputs: 0
- Stress tests: 17/17 passed
- Processing speed: approximately 12,490 samples/sec in the development environment
Architecture
V2 is a deterministic normalization system.
It does not use neural model weights and does not perform context-based spelling correction.
Intended Use
V2 can be used as:
- A Nepali NLP preprocessing layer
- A data-cleaning component
- A preprocessing stage before tokenization
- A preprocessing stage before model inference
- A safety layer around learned normalizers
Relationship to V3
V2 is an independent model/component.
It is not the same model as Supernova Nepali Normalizer V3.
V2 can optionally be used before or after V3 as a deterministic normalization and safety layer.
License
Apache-2.0 How to Use
##Supernova Nepali Normalizer V2 is a lightweight deterministic Nepali text normalization tool.
It is designed to clean common Unicode and punctuation inconsistencies in Nepali text.
Installation
Clone the repository:
git clone https://huggingface.co/Supernova11c/Supernova-Nepali-Normalizer-V2 cd Supernova-Nepali-Normalizer-V2
Python Usage
from normalizer import SupernovaNepaliNormalizer
normalizer = SupernovaNepaliNormalizer()
text = "कृपया कृपय\u200cा मलाई गणितमा कमजोर छु।"
result = normalizer.normalize(text)
print(result)
Example:
कृपया कृपया मलाई गणितमा कमजोर छु।
What V2 Does
Supernova Nepali Normalizer V2 currently provides deterministic normalization for:
- Zero-width characters
- Unicode inconsistencies
- Common punctuation variants
- Nepali text cleanup
For example:
\u200c \u200d
can be removed when they occur in unwanted positions.
Common punctuation variants are also normalized:
— → - – → - “ → " ” → " ‘ → ' ’ → '
No Neural Model Required
V2 does not require:
- GPU
- Transformers
- PyTorch
- Fine-tuned weights
- External APIs
- Internet connection during inference
It runs locally using deterministic rules.
Example
from normalizer import SupernovaNepaliNormalizer
normalizer = SupernovaNepaliNormalizer()
examples = [ "मलाई नेपाली राम्रोसँग लेख्न सिक्नुछ।", "कृपया कृपय\u200cा मलाई गणितमा कमजोर छु।", "यो एउटा—परीक्षण वाक्य हो।" ]
for text in examples: print("Input :", text) print("Output:", normalizer.normalize(text)) print()
Design Philosophy
Supernova V2 is intentionally simple and predictable.
The normalizer performs only the transformations explicitly defined by its rules. It does not generate new text or make neural predictions.
This makes V2 suitable as a preprocessing component before other NLP systems.
License
See the repository license and project files for licensing information.
⚡ Performance & CPU Benchmarks
Supernova text processing architecture is engineered for extreme, zero-overhead systems efficiency. Running entirely on standard CPU hardware without any GPU acceleration or heavy vector models, it delivers elite-tier throughput:
Language Detection & Processing: 1,237,070,359+ characters/sec
Hardware Requirement: Standard CPU (Zero GPU dependency, ultra-low memory footprint)
Architecture: Modular, deterministic, and hallucination-free text pipeline.
Test Environment: Google Colab Free Tier (Standard Shared CPU Runtime)
- Downloads last month
- 8