Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup
Paper • 2101.06983 • Published • 2
How to use wublewobble/classifier_12 with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("wublewobble/classifier_12")
sentences = [
"Event: Marriage of Plum and Jade 梅玉配\nDescription: Marriage of Plum and Jade Since its establishment in 1953, Fujian Provincial Experimental Min Opera Theatre has produced numerous classic operas, such as Marriage of Plum and Jade . For more than 53 years, Fujian Provincial Experimental Min Opera Theatre has successfully staged performances in America, Australia, Malaysia, Indonesia, Singapore, Taiwan, Hong Kong and Macao; receiving accolades and recognition from the different sources. Many overseas Chinese regard the esteemed theatre troupe as a “Cultural Messenger”. Beyond a love story, the classic Min opera Marriage of Plum and Jade also shows us the courage against the fetter of feudal ethics and the pursuit of freedom and love. One day, a scholar named Xu Jinmei visited a temple in the capital city, where he met the daughter of a governor named Su Zhenyu. They fell in love at first sight. However, Miss Su was already betrothed to Zhou Yan, a frivolous and superficial man, whom she hated unceasingly. The opera escalates to a suspenseful climax with the dramatic performance of Su’s sister-in-law, Hong Fang. She brilliantly burned a building and declared that Miss Su died accidentally in the fire. Did Miss Su perish in the fire? Will the lovers have a happy ending? The Opera profoundly portrays the innovative and vibrant regional characteristics of the East. Through a popular comedic style of performance, one can enjoy the spirit and a feast of the arts. Let’s experience the elegancy and fastidiousness of the Min Opera, and witness the astounding not- to-be-missed show.\nVenue: Esplanade Theatre",
"Theatre : Opera-Asian",
"Dance : Salsa/Tango",
"Concert : Dance Party"
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model finetuned from sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("wublewobble/classifier_12")
# Run inference
sentences = [
'Event: Shanghai Old Jazz Band 上海老爵士乐队音乐会\nDescription: Relive the golden era of Shanghai’s jazz scene with this nostalgic concert.\nVenue: Shanghai Music Hall',
'Concert : Classical Vocals-Asian',
'Dance : Modern/Contemporary',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
testBinaryClassificationEvaluator| Metric | Value |
|---|---|
| cosine_accuracy | 0.995 |
| cosine_accuracy_threshold | 0.5652 |
| cosine_f1 | 0.7417 |
| cosine_f1_threshold | 0.5227 |
| cosine_precision | 0.837 |
| cosine_recall | 0.6659 |
| cosine_ap | 0.7858 |
| cosine_mcc | 0.7442 |
anchor and positive| anchor | positive | |
|---|---|---|
| type | string | string |
| details |
|
|
| anchor | positive |
|---|---|
Event: The World of Swiss Education and Summer Camps - 2024 [G] |
Festival/Fair : Business & Professional |
Event: Wine Tasting and Sommelier Experience |
Lifestyle/Leisure : Service |
Event: Huayi 华艺节 2020 Storytellers' Wisdom - A Crosstalk Production 十五万大军直取西城而来 |
Theatre : Comedy |
CachedMultipleNegativesRankingLoss with these parameters:{
"scale": 20.0,
"similarity_fct": "cos_sim"
}
eval_strategy: stepsper_device_train_batch_size: 92per_device_eval_batch_size: 92num_train_epochs: 10warmup_ratio: 0.1batch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 92per_device_eval_batch_size: 92per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 10max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | test_cosine_ap |
|---|---|---|---|
| 0 | 0 | - | 0.1384 |
| 1.1111 | 50 | 2.0041 | 0.5899 |
| 2.2222 | 100 | 1.0715 | 0.6977 |
| 3.3333 | 150 | 0.668 | 0.7221 |
| 4.4444 | 200 | 0.4198 | 0.7442 |
| 5.5556 | 250 | 0.2544 | 0.7490 |
| 6.6667 | 300 | 0.1533 | 0.7736 |
| 7.7778 | 350 | 0.0994 | 0.7806 |
| 8.8889 | 400 | 0.066 | 0.7834 |
| 10.0 | 450 | 0.0491 | 0.7858 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}