SentenceTransformer based on BAAI/bge-m3

This is a sentence-transformers model finetuned from BAAI/bge-m3. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: BAAI/bge-m3
  • Maximum Sequence Length: 512 tokens
  • Output Dimensionality: 1024 dimensions
  • Similarity Function: Cosine Similarity

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'PeftModelForFeatureExtraction'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("DungHugging/sacombank-bge-m3-full")
# Run inference
sentences = [
    'ưu đãi Buffet hải sản giá rẻ tại Saigon Seafood',
    'hồ sơ vay tín chấp được duyệt nhanh qua App trong 1 giờ',
    'Áp dụng tiêu chuẩn CRS (Common Reporting Standard) trong trao đổi thông tin thuế quốc tế.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4611, 0.4808],
#         [0.4611, 1.0000, 0.6015],
#         [0.4808, 0.6015, 1.0000]])

Evaluation

Metrics

Semantic Similarity

Metric spec_sim final_evaluation
pearson_cosine 0.5078 0.5078
spearman_cosine 0.5003 0.5003

Binary Classification

Metric Value
cosine_accuracy 0.7399
cosine_accuracy_threshold 0.7662
cosine_f1 0.7705
cosine_f1_threshold 0.7137
cosine_precision 0.6714
cosine_recall 0.9038
cosine_ap 0.7812
cosine_mcc 0.452

Training Details

Training Dataset

Unnamed Dataset

  • Size: 2,665 training samples
  • Columns: sentence_0, sentence_1, and label
  • Approximate statistics based on the first 1000 samples:
    sentence_0 sentence_1 label
    type string string float
    details
    • min: 6 tokens
    • mean: 14.79 tokens
    • max: 30 tokens
    • min: 8 tokens
    • mean: 18.25 tokens
    • max: 49 tokens
    • min: 0.0
    • mean: 0.5
    • max: 1.0
  • Samples:
    sentence_0 sentence_1 label
    Lãi suất cơ sở (Base Rate) - dùng làm mốc tham chiếu để tính lãi vay khách hàng. Lãi suất qua đêm (Overnight Rate) - lãi suất vay nóng giữa các ngân hàng trên thị trường liên ngân hàng. 0.0
    Khoản thanh toán lớn vào cuối kỳ hạn vay. Khoản vay có cấu trúc Balloon Payment tại thời điểm đáo hạn. 1.0
    Bảo hiểm trách nhiệm dân sự chủ doanh nghiệp Bảo hiểm tai nạn con người 24/7 0.0
  • Loss: ContrastiveLoss with these parameters:
    {
        "distance_metric": "SiameseDistanceMetric.COSINE_DISTANCE",
        "margin": 0.5,
        "size_average": true
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • eval_strategy: steps
  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • num_train_epochs: 10
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: steps
  • prediction_loss_only: True
  • per_device_train_batch_size: 16
  • per_device_eval_batch_size: 16
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 5e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1
  • num_train_epochs: 10
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: {}
  • warmup_ratio: 0.0
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • bf16: False
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch_fused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss spec_sim_spearman_cosine spec_bin_cosine_ap final_evaluation_spearman_cosine
0.4970 83 - -0.1802 0.4537 -
0.9940 166 - -0.0618 0.4928 -
1.0 167 - -0.0598 0.4935 -
1.4910 249 - 0.0678 0.5431 -
1.9880 332 - 0.1480 0.5821 -
2.0 334 - 0.1500 0.5829 -
2.4850 415 - 0.2448 0.6273 -
2.9820 498 - 0.3185 0.6705 -
2.9940 500 0.0327 - - -
3.0 501 - 0.3199 0.6715 -
3.4790 581 - 0.3650 0.7018 -
3.9760 664 - 0.3993 0.7226 -
4.0 668 - 0.3986 0.7222 -
4.4731 747 - 0.4210 0.7354 -
4.9701 830 - 0.4380 0.7469 -
5.0 835 - 0.4375 0.7467 -
5.4671 913 - 0.4514 0.7525 -
5.9641 996 - 0.4606 0.7584 -
5.9880 1000 0.0241 - - -
6.0 1002 - 0.4613 0.7591 -
6.4611 1079 - 0.4717 0.7648 -
6.9581 1162 - 0.4791 0.7684 -
7.0 1169 - 0.4799 0.7689 -
7.4551 1245 - 0.4848 0.7712 -
7.9521 1328 - 0.4910 0.7750 -
8.0 1336 - 0.4915 0.7760 -
8.4491 1411 - 0.4955 0.7780 -
8.9461 1494 - 0.4972 0.7796 -
8.9820 1500 0.0217 - - -
9.0 1503 - 0.4972 0.7796 -
9.4431 1577 - 0.4995 0.7807 -
9.9401 1660 - 0.5003 0.7812 -
10.0 1670 - 0.5003 0.7812 -
-1 -1 - - - 0.5003

Framework Versions

  • Python: 3.12.12
  • Sentence Transformers: 5.1.1
  • Transformers: 4.57.1
  • PyTorch: 2.8.0+cu126
  • Accelerate: 1.11.0
  • Datasets: 4.4.2
  • Tokenizers: 0.22.1

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

ContrastiveLoss

@inproceedings{hadsell2006dimensionality,
    author={Hadsell, R. and Chopra, S. and LeCun, Y.},
    booktitle={2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06)},
    title={Dimensionality Reduction by Learning an Invariant Mapping},
    year={2006},
    volume={2},
    number={},
    pages={1735-1742},
    doi={10.1109/CVPR.2006.100}
}
Downloads last month
16
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DungHugging/sacombank-bge-m3-full

Base model

BAAI/bge-m3
Finetuned
(544)
this model

Paper for DungHugging/sacombank-bge-m3-full

Evaluation results