SentenceTransformer based on intfloat/multilingual-e5-small

This is a sentence-transformers model finetuned from intfloat/multilingual-e5-small. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: intfloat/multilingual-e5-small
  • Maximum Sequence Length: 512 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'BertModel'})
  (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("GSR-608001/avvaiyar-embedding-model")
# Run inference
sentences = [
    "I have been failing to speak no vulgarity and it's affecting me",
    '\nAathichoodi Ethical Wisdom\n\nVerse: பழிப்பன பகரேல்\n\nExplanation:\nThe refinement of a sophisticated character is reflected in the absolute refusal to utter words that are derogatory, slanderous, or condemned by the collective wisdom of society. High emotional intelligence involves recognizing the lasting impact of language on social dynamics and individual reputations, choosing instead to use speech that is constructive, dignified, and truthful. Abstaining from vulgarity and harsh rhetoric prevents the erosion of communal bonds and protects the speaker’s own reputation from being tainted by the perceived lack of self-control and empathy.\n',
    '\nAathichoodi Ethical Wisdom\n\nVerse: சக்கர நெறி நில்\n\nExplanation:\nAdherence to the established laws of the land and the universal principles of justice is fundamental to the maintenance of civilization. This guideline posits that individual actions must align with the broader legislative and ethical frameworks that govern a functioning society, often referred to as the "Dharma Chakra" or the wheel of righteousness. For modern leadership, it underscores the importance of institutional integrity and the rule of law over arbitrary power. Following this path ensures that one’s personal trajectory contributes to collective stability, fostering an environment where predictability, fairness, and systemic order empower every citizen to flourish within a structured moral ecosystem.\n',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.4474, -0.0163],
#         [ 0.4474,  1.0000,  0.0136],
#         [-0.0163,  0.0136,  1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 725 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 725 samples:
    sentence_0 sentence_1
    type string string
    details
    • min: 9 tokens
    • mean: 20.24 tokens
    • max: 45 tokens
    • min: 108 tokens
    • mean: 154.41 tokens
    • max: 181 tokens
  • Samples:
    sentence_0 sentence_1
    Paatti, teach me about don't agree with the stubborn
    Aathichoodi Ethical Wisdom

    Verse: மூர்க்கரோடு இணங்கேல்

    Explanation:
    Exercise rigorous selectivity in your social and intellectual associations by distancing yourself from the irrational and the obstinately aggressive. Engaging with those who lack emotional regulation or intellectual humility compromises your own cognitive environment and ethical standards. The influence of volatile and stubborn individuals acts as a social contagion that hinders rational discourse and collaborative evolution. Cultivating a psychological boundary against such toxicity is essential for maintaining inner stability and fostering a community grounded in reason and mutual respect.
    Paatti, teach me about don't be a glutton
    Aathichoodi Ethical Wisdom

    Verse: மீதூண் விரும்பேல்

    Explanation:
    Practice mindful consumption and biological restraint to preserve physiological harmony and cognitive clarity. Excessive intake of food or resources leads to metabolic imbalance, sensory dulling, and the erosion of self-discipline. By prioritizing nutritional efficiency over hedonistic indulgence, an individual maintains the high-functioning physical state necessary for sustained intellectual labor and ethical living. This principle extends to a modern psychological context, suggesting that an over-saturated lifestyle—whether through diet or information—stifles the pursuit of higher wisdom.
    I have been failing to read lot of books and it's affecting me
    Aathichoodi Ethical Wisdom

    Verse: நூல் பல கல்

    Explanation:
    Intellectual maturity and cognitive resilience are achieved through the systematic pursuit of diverse streams of knowledge across various disciplines. One must move beyond specialized silos to cultivate a holistic understanding of the world, utilizing literature, philosophy, and science as cognitive tools to refine critical thinking and empathy. This commitment to lifelong learning ensures that the mind remains adaptable, fostering an expansive worldview that guards against bias, ignorance, and the dangers of narrow-mindedness in a rapidly evolving global landscape.
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 32
  • num_train_epochs: 4
  • per_device_eval_batch_size: 32
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 32
  • num_train_epochs: 4
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: trackio
  • eval_strategy: no
  • per_device_eval_batch_size: 32
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Framework Versions

  • Python: 3.11.9
  • Sentence Transformers: 5.2.3
  • Transformers: 5.2.0
  • PyTorch: 2.5.1+cu121
  • Accelerate: 1.12.0
  • Datasets: 4.5.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
Downloads last month
29
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GSR-608001/avvaiyar-embedding-model

Finetuned
(185)
this model

Papers for GSR-608001/avvaiyar-embedding-model