SentenceTransformer based on sentence-transformers/stsb-distilbert-base

This is a sentence-transformers model finetuned from sentence-transformers/stsb-distilbert-base on the quora-duplicates dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'DistilBertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("kwondw/quora-mnrl")
# Run inference
queries = [
    'What are the best car gadgets in 2016?',
]
documents = [
    'What are some of the best gadgets of 2016?',
    'What is the origin of saying God Bless You after sneezing?',
    'Are there any good summer programs for high school students?',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.8092, -0.1531,  0.0052]])

Evaluation

Metrics

Information Retrieval

Metric Value
cosine_accuracy@1 0.9612
cosine_accuracy@3 0.99
cosine_accuracy@5 0.9948
cosine_accuracy@10 0.998
cosine_precision@1 0.9612
cosine_precision@3 0.4268
cosine_precision@5 0.2743
cosine_precision@10 0.1449
cosine_recall@1 0.8277
cosine_recall@3 0.9562
cosine_recall@5 0.9788
cosine_recall@10 0.9925
cosine_ndcg@10 0.9767
cosine_mrr@10 0.9761
cosine_map@100 0.968

Training Details

Training Dataset

quora-duplicates

  • Dataset: quora-duplicates at 41f6997
  • Size: 100,000 training samples
  • Columns: anchor and positive
  • Approximate statistics based on the first 100 samples:
    anchor positive
    type string string
    modality text text
    details
    • min: 6 tokens
    • mean: 14.35 tokens
    • max: 35 tokens
    • min: 7 tokens
    • mean: 14.47 tokens
    • max: 30 tokens
  • Samples:
    anchor positive
    Astrology: I am a Capricorn Sun Cap moon and cap rising...what does that say about me? I'm a triple Capricorn (Sun, Moon and ascendant in Capricorn) What does this say about me?
    How can I be a good geologist? What should I do to be a great geologist?
    How do I read and find my YouTube comments? How can I see all my Youtube comments?
  • Loss: CachedMultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "mini_batch_size": 32,
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Evaluation Dataset

quora-duplicates

  • Dataset: quora-duplicates at 41f6997
  • Size: 1,000 evaluation samples
  • Columns: anchor and positive
  • Approximate statistics based on the first 100 samples:
    anchor positive
    type string string
    modality text text
    details
    • min: 6 tokens
    • mean: 14.36 tokens
    • max: 38 tokens
    • min: 8 tokens
    • mean: 14.38 tokens
    • max: 38 tokens
  • Samples:
    anchor positive
    What is the best English translation of the Bhagavad Gita? Which is the best English version of Bhagavad-Gita?
    Quora kept refreshing on its own. Is this a normal thing or is it just me? Why does Quora keep refreshing the page?
    What is it like to study in McGill University? What is it like to study at McGill University?
  • Loss: CachedMultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "mini_batch_size": 32,
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 64
  • num_train_epochs: 1
  • learning_rate: 2e-05
  • warmup_steps: 0.1
  • fp16: True
  • per_device_eval_batch_size: 64
  • load_best_model_at_end: True
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 64
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: True
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 64
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss Validation Loss quora-ir_cosine_ndcg@10
-1 -1 - - 0.9415
0.0640 100 0.1075 - -
0.1280 200 0.0788 - -
0.1599 250 - 0.0425 0.9656
0.1919 300 0.0625 - -
0.2559 400 0.0601 - -
0.3199 500 0.0626 0.0383 0.9702
0.3839 600 0.0513 - -
0.4479 700 0.0450 - -
0.4798 750 - 0.0383 0.9714
0.5118 800 0.0476 - -
0.5758 900 0.0514 - -
0.6398 1000 0.0383 0.0365 0.9735
0.7038 1100 0.0488 - -
0.7678 1200 0.0425 - -
0.7997 1250 - 0.0344 0.9742
0.8317 1300 0.0495 - -
0.8957 1400 0.0379 - -
0.9597 1500 0.0481 0.0340 0.9749
1.0 1563 - 0.0341 0.9748
-1 -1 - - 0.9749
0.0640 100 0.0408 - -
0.1280 200 0.0322 - -
0.1599 250 - 0.0342 0.9727
0.1919 300 0.0277 - -
0.2559 400 0.0294 - -
0.3199 500 0.0324 0.0365 0.9742
0.3839 600 0.0327 - -
0.4479 700 0.0284 - -
0.4798 750 - 0.0335 0.9750
0.5118 800 0.0300 - -
0.5758 900 0.0360 - -
0.6398 1000 0.0261 0.0350 0.9753
0.7038 1100 0.0353 - -
0.7678 1200 0.0335 - -
0.7997 1250 - 0.0331 0.9763
0.8317 1300 0.0388 - -
0.8957 1400 0.0308 - -
0.9597 1500 0.0410 0.0327 0.9767
1.0 1563 - 0.0328 0.9767
-1 -1 - - 0.9767
  • The bold row denotes the saved checkpoint.

Training Time

  • Training: 10.7 minutes

Framework Versions

  • Python: 3.12.13
  • Sentence Transformers: 5.6.0
  • Transformers: 5.13.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.14.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

CachedMultipleNegativesRankingLoss

@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
Downloads last month
34
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kwondw/quora-mnrl

Finetuned
(10)
this model

Dataset used to train kwondw/quora-mnrl

Papers for kwondw/quora-mnrl

Evaluation results