Feature Extraction
sentence-transformers
Safetensors
Hindi
English
bert
hindi
hinglish
code-mixed
song-lyrics
tsdae
text-embeddings-inference
Instructions to use meet5568/bhash_finetune with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use meet5568/bhash_finetune with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("meet5568/bhash_finetune") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
| language: | |
| - hi | |
| - en | |
| license: apache-2.0 | |
| library_name: sentence-transformers | |
| tags: | |
| - sentence-transformers | |
| - feature-extraction | |
| - hindi | |
| - hinglish | |
| - code-mixed | |
| - song-lyrics | |
| - tsdae | |
| base_model: AkshitaS/bhasha-embed-v0 | |
| # BhashaEmbed Hindi Songs — TSDAE Domain Adapted | |
| Domain-adapted version of [`AkshitaS/bhasha-embed-v0`](https://huggingface.co/AkshitaS/bhasha-embed-v0) | |
| for Hindi film song lyrics. | |
| ## Training | |
| - **Method:** TSDAE full fine-tuning (no quantization, no LoRA) | |
| - **Fine-tuning:** Full parameter fine-tuning (all weights updated) | |
| - **Corpus:** 8,464 Hindi film songs, 28,713 verse-level chunks | |
| - **Scripts:** Devanagari, Romanized Hindi (Hinglish), English — single unified model | |
| - **Script balancing:** weighted sampling to 44/44/12 (Devanagari/Hinglish/English) | |
| - **Epochs:** 5 | |
| - **Batch size:** 4 | |
| - **Learning rate:** 3e-05 | |
| - **Max sequence length:** 400 subword tokens | |
| ## Intended Use | |
| Sentence embeddings for Hindi film song lyrics across all three script forms. | |
| Designed for thematic classification and generational/temporal analysis of Hindi music. | |
| ## Limitations | |
| - Full parameter fine-tuning — no LoRA/quantization constraints on how far | |
| weights can move from the base model | |
| - No supervised thematic fine-tuning applied (Phase A only) | |
| - Not evaluated on tasks outside Hindi song lyrics | |
| ## Usage | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| model = SentenceTransformer("meet5568/bhash_finetune") | |
| lyrics = [ | |
| "तेरे बिना जिंदगी से कोई शिकवा नहीं", | |
| "tere bina zindagi se koi shikwa nahi", | |
| "without you life has no complaint", | |
| ] | |
| embeddings = model.encode(lyrics) | |
| print(embeddings.shape) # (3, 768) | |
| ``` | |