Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Paper • 1908.10084 • Published • 17
How to use DungHugging/sacombank-bge-m3-full with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("DungHugging/sacombank-bge-m3-full")
sentences = [
"Cơ chế một cửa quốc gia về xuất nhập khẩu.",
"Kết nối cổng National Single Window để thanh toán thuế hải quan.",
"Thước đo mức độ suy giảm giá trị của quyền chọn (Time Decay) khi thời gian trôi dần về ngày đáo hạn.",
"Gửi tiền không kỳ hạn cho doanh nghiệp"
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model finetuned from BAAI/bge-m3. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'PeftModelForFeatureExtraction'})
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("DungHugging/sacombank-bge-m3-full")
# Run inference
sentences = [
'ưu đãi Buffet hải sản giá rẻ tại Saigon Seafood',
'hồ sơ vay tín chấp được duyệt nhanh qua App trong 1 giờ',
'Áp dụng tiêu chuẩn CRS (Common Reporting Standard) trong trao đổi thông tin thuế quốc tế.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4611, 0.4808],
# [0.4611, 1.0000, 0.6015],
# [0.4808, 0.6015, 1.0000]])
spec_sim and final_evaluationEmbeddingSimilarityEvaluator| Metric | spec_sim | final_evaluation |
|---|---|---|
| pearson_cosine | 0.5078 | 0.5078 |
| spearman_cosine | 0.5003 | 0.5003 |
spec_binBinaryClassificationEvaluator| Metric | Value |
|---|---|
| cosine_accuracy | 0.7399 |
| cosine_accuracy_threshold | 0.7662 |
| cosine_f1 | 0.7705 |
| cosine_f1_threshold | 0.7137 |
| cosine_precision | 0.6714 |
| cosine_recall | 0.9038 |
| cosine_ap | 0.7812 |
| cosine_mcc | 0.452 |
sentence_0, sentence_1, and label| sentence_0 | sentence_1 | label | |
|---|---|---|---|
| type | string | string | float |
| details |
|
|
|
| sentence_0 | sentence_1 | label |
|---|---|---|
Lãi suất cơ sở (Base Rate) - dùng làm mốc tham chiếu để tính lãi vay khách hàng. |
Lãi suất qua đêm (Overnight Rate) - lãi suất vay nóng giữa các ngân hàng trên thị trường liên ngân hàng. |
0.0 |
Khoản thanh toán lớn vào cuối kỳ hạn vay. |
Khoản vay có cấu trúc Balloon Payment tại thời điểm đáo hạn. |
1.0 |
Bảo hiểm trách nhiệm dân sự chủ doanh nghiệp |
Bảo hiểm tai nạn con người 24/7 |
0.0 |
ContrastiveLoss with these parameters:{
"distance_metric": "SiameseDistanceMetric.COSINE_DISTANCE",
"margin": 0.5,
"size_average": true
}
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16num_train_epochs: 10multi_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 10max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torch_fusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthproject: huggingfacetrackio_space_id: trackioddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: noneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Trueprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robinrouter_mapping: {}learning_rate_mapping: {}| Epoch | Step | Training Loss | spec_sim_spearman_cosine | spec_bin_cosine_ap | final_evaluation_spearman_cosine |
|---|---|---|---|---|---|
| 0.4970 | 83 | - | -0.1802 | 0.4537 | - |
| 0.9940 | 166 | - | -0.0618 | 0.4928 | - |
| 1.0 | 167 | - | -0.0598 | 0.4935 | - |
| 1.4910 | 249 | - | 0.0678 | 0.5431 | - |
| 1.9880 | 332 | - | 0.1480 | 0.5821 | - |
| 2.0 | 334 | - | 0.1500 | 0.5829 | - |
| 2.4850 | 415 | - | 0.2448 | 0.6273 | - |
| 2.9820 | 498 | - | 0.3185 | 0.6705 | - |
| 2.9940 | 500 | 0.0327 | - | - | - |
| 3.0 | 501 | - | 0.3199 | 0.6715 | - |
| 3.4790 | 581 | - | 0.3650 | 0.7018 | - |
| 3.9760 | 664 | - | 0.3993 | 0.7226 | - |
| 4.0 | 668 | - | 0.3986 | 0.7222 | - |
| 4.4731 | 747 | - | 0.4210 | 0.7354 | - |
| 4.9701 | 830 | - | 0.4380 | 0.7469 | - |
| 5.0 | 835 | - | 0.4375 | 0.7467 | - |
| 5.4671 | 913 | - | 0.4514 | 0.7525 | - |
| 5.9641 | 996 | - | 0.4606 | 0.7584 | - |
| 5.9880 | 1000 | 0.0241 | - | - | - |
| 6.0 | 1002 | - | 0.4613 | 0.7591 | - |
| 6.4611 | 1079 | - | 0.4717 | 0.7648 | - |
| 6.9581 | 1162 | - | 0.4791 | 0.7684 | - |
| 7.0 | 1169 | - | 0.4799 | 0.7689 | - |
| 7.4551 | 1245 | - | 0.4848 | 0.7712 | - |
| 7.9521 | 1328 | - | 0.4910 | 0.7750 | - |
| 8.0 | 1336 | - | 0.4915 | 0.7760 | - |
| 8.4491 | 1411 | - | 0.4955 | 0.7780 | - |
| 8.9461 | 1494 | - | 0.4972 | 0.7796 | - |
| 8.9820 | 1500 | 0.0217 | - | - | - |
| 9.0 | 1503 | - | 0.4972 | 0.7796 | - |
| 9.4431 | 1577 | - | 0.4995 | 0.7807 | - |
| 9.9401 | 1660 | - | 0.5003 | 0.7812 | - |
| 10.0 | 1670 | - | 0.5003 | 0.7812 | - |
| -1 | -1 | - | - | - | 0.5003 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@inproceedings{hadsell2006dimensionality,
author={Hadsell, R. and Chopra, S. and LeCun, Y.},
booktitle={2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06)},
title={Dimensionality Reduction by Learning an Invariant Mapping},
year={2006},
volume={2},
number={},
pages={1735-1742},
doi={10.1109/CVPR.2006.100}
}
Base model
BAAI/bge-m3