wublewobble commited on
Commit
9a1da67
·
verified ·
1 Parent(s): 3242097

Add new SentenceTransformer model

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ unigram.json filter=lfs diff=lfs merge=lfs -text
1_Pooling/config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "word_embedding_dimension": 384,
3
+ "pooling_mode_cls_token": false,
4
+ "pooling_mode_mean_tokens": true,
5
+ "pooling_mode_max_tokens": false,
6
+ "pooling_mode_mean_sqrt_len_tokens": false,
7
+ "pooling_mode_weightedmean_tokens": false,
8
+ "pooling_mode_lasttoken": false,
9
+ "include_prompt": true
10
+ }
README.md ADDED
@@ -0,0 +1,506 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - sentence-transformers
4
+ - sentence-similarity
5
+ - feature-extraction
6
+ - generated_from_trainer
7
+ - dataset_size:4140
8
+ - loss:CachedMultipleNegativesRankingLoss
9
+ base_model: sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
10
+ widget:
11
+ - source_sentence: 'Event: Marriage of Plum and Jade 梅玉配
12
+
13
+ Description: Marriage of Plum and Jade   Since its establishment in 1953, Fujian
14
+ Provincial Experimental Min Opera Theatre has produced numerous classic operas,
15
+ such as Marriage of Plum and Jade . For more than 53 years, Fujian Provincial
16
+ Experimental Min Opera Theatre has successfully staged performances in America,
17
+ Australia, Malaysia, Indonesia, Singapore, Taiwan, Hong Kong and Macao; receiving
18
+ accolades and recognition from the different sources. Many overseas Chinese regard
19
+ the esteemed theatre troupe as a “Cultural Messenger”.   Beyond a love story,
20
+ the classic Min opera Marriage of Plum and Jade also shows us the courage against
21
+ the fetter of feudal ethics and the pursuit of freedom and love. One day, a scholar
22
+ named Xu Jinmei visited a temple in the capital city, where he met the daughter
23
+ of a governor named Su Zhenyu. They fell in love at first sight. However, Miss
24
+ Su was already betrothed to Zhou Yan, a frivolous and superficial man, whom she
25
+ hated unceasingly.  The opera escalates to a suspenseful climax with the dramatic
26
+ performance of Su’s sister-in-law, Hong Fang. She brilliantly burned a building
27
+ and declared that Miss Su died accidentally in the fire. Did Miss Su perish in
28
+ the fire? Will the lovers have a happy ending?   The Opera profoundly portrays
29
+ the innovative and vibrant regional characteristics of the East. Through a popular
30
+ comedic style of performance, one can enjoy the spirit and a feast of the arts.
31
+ Let’s experience the elegancy and fastidiousness of the Min Opera, and witness
32
+ the astounding not- to-be-missed show.
33
+
34
+ Venue: Esplanade Theatre'
35
+ sentences:
36
+ - 'Theatre : Opera-Asian'
37
+ - 'Dance : Salsa/Tango'
38
+ - 'Concert : Dance Party'
39
+ - source_sentence: 'Event: Street Art and Culture Festival [G]
40
+
41
+ Description: This festival brings together contemporary street art, live graffiti
42
+ performances, and culture-focused exhibitions to celebrate urban creativity, inspiring
43
+ the public to appreciate the street art movement.
44
+
45
+ Venue: Various Streets in Singapore'
46
+ sentences:
47
+ - 'Dance : Traditional/Ethnic'
48
+ - 'Festival/Fair : Community & Culture'
49
+ - 'Seminar/Workshop : Corporate'
50
+ - source_sentence: 'Event: Cultivating a Collaborative Marriage - For Better & Forever
51
+ - 9.30am (Recommended for all couples)
52
+
53
+ Description: Marriage Convention 2015 Growing Together. Staying Together. Marriage
54
+ is like a seed of love sown by two people, and its resulting growth is a representation
55
+ of the effort put into nurturing it. Just like the saying, “Marriage is a journey,
56
+ not a destination.” - it is an ongoing process that requires both parties to commit
57
+ and keep working on it, or run the risk of drifting apart. Marriage Convention
58
+ 2015 brings to you “ Growing Together. Staying Together” . Hear from overseas
59
+ and local marriage experts who will share insights into how couples can build,
60
+ grow together and strengthen their marriage. Get tips and ideas from renowned
61
+ author, clinician in marriage and family therapy, Dr Sherod Miller and his wife,
62
+ Dr Phyllis, on how to achieve a thriving marriage. Marriage Convention 2015
63
+ is brought to you by Families for Life.  To find out more, please visit www.toggle.sg/marriageconvention
64
+
65
+ Venue: Suntec Singapore Convention and Exhibition Centre, Summit 2, Level 3'
66
+ sentences:
67
+ - 'Concert : Pop-Western'
68
+ - 'Theatre : Children'
69
+ - 'Seminar/Workshop : Family & Social'
70
+ - source_sentence: 'Event: Night at the Movies
71
+
72
+ Description: Alex Tan Sing is undoubtedly Singapore’s most versatile, most creative
73
+ multilingual, multi-talented artiste. www.AlexTanSing.com Born & bred Singaporean,
74
+ being a Singer-Songwriter-Entertainer-Comedian-Showhost-Show Producer for well
75
+ over 20 years & now a Content Creator-Live Streamer-Edutainment Provider with
76
+ the advent of the pandemic, he returns to his very first loves --- singing & entertaining. Join
77
+ Alex for FREE as he brightens up your Saturday Night at 9pm with an hour-long
78
+ live concerts-with his music, singing, comedic impressions & witty interactions.
79
+ His entertaining Live Stream on Zoom is his way of giving back as he uplifts your
80
+ spirits & moves you with his passion during these difficult times & his genuine
81
+ love & compassion for the community. Expect vibrant uplifting originals such
82
+ as A Happy Tune, Make You Smile, The Rainbow Song, Sweet Sweet Love & Sunrise to
83
+ cover versions of songs from various genres as he creates the best online experience
84
+ & parties, together with his special guests! Catch the new release of his Music
85
+ Video against Covid-19--- Stop The Virus is a fun & easy song to educate us
86
+ all in our fight against the virus!  Join him on Zoom absolutely FREE & come
87
+ dressed every Saturday night in the various themes for the best online party experience
88
+ & stand to win shopping vouchers for the Best Dressed & have fun with his interactive
89
+ games too! Come dressed to the various themes:-  29 Aug 2020    Night at the
90
+ Movies 05 Sep 2020    Disney Magic 12 Sep 2020    Saturday Night Fever 19 Sep
91
+ 2020    Rock & Roll 26 Sep 2020    Back to the 80''s-Solid Gold Alex will attempt
92
+ to transport you through time & space to another era, another world with his special
93
+ renditions of perennial favourites, Top 40’s, Evergreens, Disco grooves & moves,
94
+ 80’s Retro, from Broadway’s best to Disney Delights too! It will be the best
95
+ times of your lives as Alex Tan Sing creates the most memorable & entertaining
96
+ online experience!!!  Fun & Free!!!
97
+
98
+ Venue: Zoom Space'
99
+ sentences:
100
+ - 'Concert : Jazz'
101
+ - 'Musical : Western'
102
+ - 'Food & Beverage : F & B Voucher'
103
+ - source_sentence: 'Event: Shanghai Old Jazz Band 上海老爵士乐队音乐会
104
+
105
+ Description: Relive the golden era of Shanghai’s jazz scene with this nostalgic
106
+ concert.
107
+
108
+ Venue: Shanghai Music Hall'
109
+ sentences:
110
+ - 'Sports : Chess'
111
+ - 'Dance : Modern/Contemporary'
112
+ - 'Concert : Classical Vocals-Asian'
113
+ pipeline_tag: sentence-similarity
114
+ library_name: sentence-transformers
115
+ metrics:
116
+ - cosine_accuracy
117
+ - cosine_accuracy_threshold
118
+ - cosine_f1
119
+ - cosine_f1_threshold
120
+ - cosine_precision
121
+ - cosine_recall
122
+ - cosine_ap
123
+ - cosine_mcc
124
+ model-index:
125
+ - name: SentenceTransformer based on sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
126
+ results:
127
+ - task:
128
+ type: binary-classification
129
+ name: Binary Classification
130
+ dataset:
131
+ name: test
132
+ type: test
133
+ metrics:
134
+ - type: cosine_accuracy
135
+ value: 0.9950316493490983
136
+ name: Cosine Accuracy
137
+ - type: cosine_accuracy_threshold
138
+ value: 0.5651620030403137
139
+ name: Cosine Accuracy Threshold
140
+ - type: cosine_f1
141
+ value: 0.7417380660954712
142
+ name: Cosine F1
143
+ - type: cosine_f1_threshold
144
+ value: 0.5227307081222534
145
+ name: Cosine F1 Threshold
146
+ - type: cosine_precision
147
+ value: 0.8370165745856354
148
+ name: Cosine Precision
149
+ - type: cosine_recall
150
+ value: 0.6659340659340659
151
+ name: Cosine Recall
152
+ - type: cosine_ap
153
+ value: 0.7858178723117079
154
+ name: Cosine Ap
155
+ - type: cosine_mcc
156
+ value: 0.7441583162870657
157
+ name: Cosine Mcc
158
+ ---
159
+
160
+ # SentenceTransformer based on sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
161
+
162
+ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2). It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
163
+
164
+ ## Model Details
165
+
166
+ ### Model Description
167
+ - **Model Type:** Sentence Transformer
168
+ - **Base model:** [sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2) <!-- at revision 86741b4e3f5cb7765a600d3a3d55a0f6a6cb443d -->
169
+ - **Maximum Sequence Length:** 128 tokens
170
+ - **Output Dimensionality:** 384 dimensions
171
+ - **Similarity Function:** Cosine Similarity
172
+ <!-- - **Training Dataset:** Unknown -->
173
+ <!-- - **Language:** Unknown -->
174
+ <!-- - **License:** Unknown -->
175
+
176
+ ### Model Sources
177
+
178
+ - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
179
+ - **Repository:** [Sentence Transformers on GitHub](https://github.com/UKPLab/sentence-transformers)
180
+ - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)
181
+
182
+ ### Full Model Architecture
183
+
184
+ ```
185
+ SentenceTransformer(
186
+ (0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: BertModel
187
+ (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
188
+ )
189
+ ```
190
+
191
+ ## Usage
192
+
193
+ ### Direct Usage (Sentence Transformers)
194
+
195
+ First install the Sentence Transformers library:
196
+
197
+ ```bash
198
+ pip install -U sentence-transformers
199
+ ```
200
+
201
+ Then you can load this model and run inference.
202
+ ```python
203
+ from sentence_transformers import SentenceTransformer
204
+
205
+ # Download from the 🤗 Hub
206
+ model = SentenceTransformer("wublewobble/classifier_12")
207
+ # Run inference
208
+ sentences = [
209
+ 'Event: Shanghai Old Jazz Band 上海老爵士乐队音乐会\nDescription: Relive the golden era of Shanghai’s jazz scene with this nostalgic concert.\nVenue: Shanghai Music Hall',
210
+ 'Concert : Classical Vocals-Asian',
211
+ 'Dance : Modern/Contemporary',
212
+ ]
213
+ embeddings = model.encode(sentences)
214
+ print(embeddings.shape)
215
+ # [3, 384]
216
+
217
+ # Get the similarity scores for the embeddings
218
+ similarities = model.similarity(embeddings, embeddings)
219
+ print(similarities.shape)
220
+ # [3, 3]
221
+ ```
222
+
223
+ <!--
224
+ ### Direct Usage (Transformers)
225
+
226
+ <details><summary>Click to see the direct usage in Transformers</summary>
227
+
228
+ </details>
229
+ -->
230
+
231
+ <!--
232
+ ### Downstream Usage (Sentence Transformers)
233
+
234
+ You can finetune this model on your own dataset.
235
+
236
+ <details><summary>Click to expand</summary>
237
+
238
+ </details>
239
+ -->
240
+
241
+ <!--
242
+ ### Out-of-Scope Use
243
+
244
+ *List how the model may foreseeably be misused and address what users ought not to do with the model.*
245
+ -->
246
+
247
+ ## Evaluation
248
+
249
+ ### Metrics
250
+
251
+ #### Binary Classification
252
+
253
+ * Dataset: `test`
254
+ * Evaluated with [<code>BinaryClassificationEvaluator</code>](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.BinaryClassificationEvaluator)
255
+
256
+ | Metric | Value |
257
+ |:--------------------------|:-----------|
258
+ | cosine_accuracy | 0.995 |
259
+ | cosine_accuracy_threshold | 0.5652 |
260
+ | cosine_f1 | 0.7417 |
261
+ | cosine_f1_threshold | 0.5227 |
262
+ | cosine_precision | 0.837 |
263
+ | cosine_recall | 0.6659 |
264
+ | **cosine_ap** | **0.7858** |
265
+ | cosine_mcc | 0.7442 |
266
+
267
+ <!--
268
+ ## Bias, Risks and Limitations
269
+
270
+ *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
271
+ -->
272
+
273
+ <!--
274
+ ### Recommendations
275
+
276
+ *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
277
+ -->
278
+
279
+ ## Training Details
280
+
281
+ ### Training Dataset
282
+
283
+ #### Unnamed Dataset
284
+
285
+ * Size: 4,140 training samples
286
+ * Columns: <code>anchor</code> and <code>positive</code>
287
+ * Approximate statistics based on the first 1000 samples:
288
+ | | anchor | positive |
289
+ |:--------|:------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------|
290
+ | type | string | string |
291
+ | details | <ul><li>min: 21 tokens</li><li>mean: 99.72 tokens</li><li>max: 128 tokens</li></ul> | <ul><li>min: 5 tokens</li><li>mean: 8.29 tokens</li><li>max: 21 tokens</li></ul> |
292
+ * Samples:
293
+ | anchor | positive |
294
+ |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------|
295
+ | <code>Event: The World of Swiss Education and Summer Camps - 2024 [G]<br>Description: Come with your children to discover Switzerland's most esteemed boarding schools, hotel management schools and summer camps on a fun-filled family adventure through the Alps, experience Swiss culture and meet admission directors in person. 3.00pm Doors open. Families are free to discover boarding schools, summer camps & network. Children’s activities begin. 3.10pm Welcome by the Ambassador of Switzerland, HE Frank Grütter. Presentations to introduce Swiss schools and summer camps. 3.30pm Lucky draw 5.00pm Close<br>Venue: The Embassy Room, St. Regis Hotel</code> | <code>Festival/Fair : Business & Professional</code> |
296
+ | <code>Event: Wine Tasting and Sommelier Experience<br>Description: Join our expert sommelier for an immersive wine tasting experience. Sample premium wines, learn the art of wine pairing, and develop a deeper appreciation for fine wines in an intimate setting.<br>Venue: Wine Tasting Room</code> | <code>Lifestyle/Leisure : Service</code> |
297
+ | <code>Event: Huayi 华艺节 2020 Storytellers' Wisdom - A Crosstalk Production 十五万大军直取西城而来<br>Description: What does Detective Conan and Justice Bao have in common? Can the capable Sun Wukong with his endless transformations survive in the modern society? What can Jin Yong’s stories tell you about the philosophy of ‘three’? With a focus on the Empty Fort Strategy from the classic Chinese military directives Thirty-Six Stratagems , Storytellers’ Wisdom is a lighthearted crosstalk production that enacts the various chapters of Chinese culture and history through an engaging performance filled with clever dialogue and witty humour. An original creation by the renowned Comedians Workshop from Taiwan, Storytellers’ Wisdom features a selection of the group’s best works performed by established theatre practitioners including Feng Yi-Gang and Sung Shao-Ching. Discover humorous anecdotes about life told through stories from Romance of the Three Kingdoms , Justice Bao, Sun Wukong and classics fro...</code> | <code>Theatre : Comedy</code> |
298
+ * Loss: [<code>CachedMultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters:
299
+ ```json
300
+ {
301
+ "scale": 20.0,
302
+ "similarity_fct": "cos_sim"
303
+ }
304
+ ```
305
+
306
+ ### Training Hyperparameters
307
+ #### Non-Default Hyperparameters
308
+
309
+ - `eval_strategy`: steps
310
+ - `per_device_train_batch_size`: 92
311
+ - `per_device_eval_batch_size`: 92
312
+ - `num_train_epochs`: 10
313
+ - `warmup_ratio`: 0.1
314
+ - `batch_sampler`: no_duplicates
315
+
316
+ #### All Hyperparameters
317
+ <details><summary>Click to expand</summary>
318
+
319
+ - `overwrite_output_dir`: False
320
+ - `do_predict`: False
321
+ - `eval_strategy`: steps
322
+ - `prediction_loss_only`: True
323
+ - `per_device_train_batch_size`: 92
324
+ - `per_device_eval_batch_size`: 92
325
+ - `per_gpu_train_batch_size`: None
326
+ - `per_gpu_eval_batch_size`: None
327
+ - `gradient_accumulation_steps`: 1
328
+ - `eval_accumulation_steps`: None
329
+ - `torch_empty_cache_steps`: None
330
+ - `learning_rate`: 5e-05
331
+ - `weight_decay`: 0.0
332
+ - `adam_beta1`: 0.9
333
+ - `adam_beta2`: 0.999
334
+ - `adam_epsilon`: 1e-08
335
+ - `max_grad_norm`: 1.0
336
+ - `num_train_epochs`: 10
337
+ - `max_steps`: -1
338
+ - `lr_scheduler_type`: linear
339
+ - `lr_scheduler_kwargs`: {}
340
+ - `warmup_ratio`: 0.1
341
+ - `warmup_steps`: 0
342
+ - `log_level`: passive
343
+ - `log_level_replica`: warning
344
+ - `log_on_each_node`: True
345
+ - `logging_nan_inf_filter`: True
346
+ - `save_safetensors`: True
347
+ - `save_on_each_node`: False
348
+ - `save_only_model`: False
349
+ - `restore_callback_states_from_checkpoint`: False
350
+ - `no_cuda`: False
351
+ - `use_cpu`: False
352
+ - `use_mps_device`: False
353
+ - `seed`: 42
354
+ - `data_seed`: None
355
+ - `jit_mode_eval`: False
356
+ - `use_ipex`: False
357
+ - `bf16`: False
358
+ - `fp16`: False
359
+ - `fp16_opt_level`: O1
360
+ - `half_precision_backend`: auto
361
+ - `bf16_full_eval`: False
362
+ - `fp16_full_eval`: False
363
+ - `tf32`: None
364
+ - `local_rank`: 0
365
+ - `ddp_backend`: None
366
+ - `tpu_num_cores`: None
367
+ - `tpu_metrics_debug`: False
368
+ - `debug`: []
369
+ - `dataloader_drop_last`: False
370
+ - `dataloader_num_workers`: 0
371
+ - `dataloader_prefetch_factor`: None
372
+ - `past_index`: -1
373
+ - `disable_tqdm`: False
374
+ - `remove_unused_columns`: True
375
+ - `label_names`: None
376
+ - `load_best_model_at_end`: False
377
+ - `ignore_data_skip`: False
378
+ - `fsdp`: []
379
+ - `fsdp_min_num_params`: 0
380
+ - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
381
+ - `fsdp_transformer_layer_cls_to_wrap`: None
382
+ - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
383
+ - `deepspeed`: None
384
+ - `label_smoothing_factor`: 0.0
385
+ - `optim`: adamw_torch
386
+ - `optim_args`: None
387
+ - `adafactor`: False
388
+ - `group_by_length`: False
389
+ - `length_column_name`: length
390
+ - `ddp_find_unused_parameters`: None
391
+ - `ddp_bucket_cap_mb`: None
392
+ - `ddp_broadcast_buffers`: False
393
+ - `dataloader_pin_memory`: True
394
+ - `dataloader_persistent_workers`: False
395
+ - `skip_memory_metrics`: True
396
+ - `use_legacy_prediction_loop`: False
397
+ - `push_to_hub`: False
398
+ - `resume_from_checkpoint`: None
399
+ - `hub_model_id`: None
400
+ - `hub_strategy`: every_save
401
+ - `hub_private_repo`: None
402
+ - `hub_always_push`: False
403
+ - `gradient_checkpointing`: False
404
+ - `gradient_checkpointing_kwargs`: None
405
+ - `include_inputs_for_metrics`: False
406
+ - `include_for_metrics`: []
407
+ - `eval_do_concat_batches`: True
408
+ - `fp16_backend`: auto
409
+ - `push_to_hub_model_id`: None
410
+ - `push_to_hub_organization`: None
411
+ - `mp_parameters`:
412
+ - `auto_find_batch_size`: False
413
+ - `full_determinism`: False
414
+ - `torchdynamo`: None
415
+ - `ray_scope`: last
416
+ - `ddp_timeout`: 1800
417
+ - `torch_compile`: False
418
+ - `torch_compile_backend`: None
419
+ - `torch_compile_mode`: None
420
+ - `dispatch_batches`: None
421
+ - `split_batches`: None
422
+ - `include_tokens_per_second`: False
423
+ - `include_num_input_tokens_seen`: False
424
+ - `neftune_noise_alpha`: None
425
+ - `optim_target_modules`: None
426
+ - `batch_eval_metrics`: False
427
+ - `eval_on_start`: False
428
+ - `use_liger_kernel`: False
429
+ - `eval_use_gather_object`: False
430
+ - `average_tokens_across_devices`: False
431
+ - `prompts`: None
432
+ - `batch_sampler`: no_duplicates
433
+ - `multi_dataset_batch_sampler`: proportional
434
+
435
+ </details>
436
+
437
+ ### Training Logs
438
+ | Epoch | Step | Training Loss | test_cosine_ap |
439
+ |:------:|:----:|:-------------:|:--------------:|
440
+ | 0 | 0 | - | 0.1384 |
441
+ | 1.1111 | 50 | 2.0041 | 0.5899 |
442
+ | 2.2222 | 100 | 1.0715 | 0.6977 |
443
+ | 3.3333 | 150 | 0.668 | 0.7221 |
444
+ | 4.4444 | 200 | 0.4198 | 0.7442 |
445
+ | 5.5556 | 250 | 0.2544 | 0.7490 |
446
+ | 6.6667 | 300 | 0.1533 | 0.7736 |
447
+ | 7.7778 | 350 | 0.0994 | 0.7806 |
448
+ | 8.8889 | 400 | 0.066 | 0.7834 |
449
+ | 10.0 | 450 | 0.0491 | 0.7858 |
450
+
451
+
452
+ ### Framework Versions
453
+ - Python: 3.11.11
454
+ - Sentence Transformers: 3.4.1
455
+ - Transformers: 4.48.3
456
+ - PyTorch: 2.6.0+cu124
457
+ - Accelerate: 1.3.0
458
+ - Datasets: 3.4.1
459
+ - Tokenizers: 0.21.1
460
+
461
+ ## Citation
462
+
463
+ ### BibTeX
464
+
465
+ #### Sentence Transformers
466
+ ```bibtex
467
+ @inproceedings{reimers-2019-sentence-bert,
468
+ title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
469
+ author = "Reimers, Nils and Gurevych, Iryna",
470
+ booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
471
+ month = "11",
472
+ year = "2019",
473
+ publisher = "Association for Computational Linguistics",
474
+ url = "https://arxiv.org/abs/1908.10084",
475
+ }
476
+ ```
477
+
478
+ #### CachedMultipleNegativesRankingLoss
479
+ ```bibtex
480
+ @misc{gao2021scaling,
481
+ title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
482
+ author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
483
+ year={2021},
484
+ eprint={2101.06983},
485
+ archivePrefix={arXiv},
486
+ primaryClass={cs.LG}
487
+ }
488
+ ```
489
+
490
+ <!--
491
+ ## Glossary
492
+
493
+ *Clearly define terms in order to be accessible across audiences.*
494
+ -->
495
+
496
+ <!--
497
+ ## Model Card Authors
498
+
499
+ *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
500
+ -->
501
+
502
+ <!--
503
+ ## Model Card Contact
504
+
505
+ *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
506
+ -->
config.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2",
3
+ "architectures": [
4
+ "BertModel"
5
+ ],
6
+ "attention_probs_dropout_prob": 0.1,
7
+ "classifier_dropout": null,
8
+ "gradient_checkpointing": false,
9
+ "hidden_act": "gelu",
10
+ "hidden_dropout_prob": 0.1,
11
+ "hidden_size": 384,
12
+ "initializer_range": 0.02,
13
+ "intermediate_size": 1536,
14
+ "layer_norm_eps": 1e-12,
15
+ "max_position_embeddings": 512,
16
+ "model_type": "bert",
17
+ "num_attention_heads": 12,
18
+ "num_hidden_layers": 12,
19
+ "pad_token_id": 0,
20
+ "position_embedding_type": "absolute",
21
+ "torch_dtype": "float32",
22
+ "transformers_version": "4.48.3",
23
+ "type_vocab_size": 2,
24
+ "use_cache": true,
25
+ "vocab_size": 250037
26
+ }
config_sentence_transformers.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "__version__": {
3
+ "sentence_transformers": "3.4.1",
4
+ "transformers": "4.48.3",
5
+ "pytorch": "2.6.0+cu124"
6
+ },
7
+ "prompts": {},
8
+ "default_prompt_name": null,
9
+ "similarity_fn_name": "cosine"
10
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0a42337e67e5200dcb331f51d6f0dfcec94de1a34c4dbd77fd77509962bcac96
3
+ size 470637416
modules.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "idx": 0,
4
+ "name": "0",
5
+ "path": "",
6
+ "type": "sentence_transformers.models.Transformer"
7
+ },
8
+ {
9
+ "idx": 1,
10
+ "name": "1",
11
+ "path": "1_Pooling",
12
+ "type": "sentence_transformers.models.Pooling"
13
+ }
14
+ ]
sentence_bert_config.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "max_seq_length": 128,
3
+ "do_lower_case": false
4
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "cls_token": {
10
+ "content": "<s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "eos_token": {
17
+ "content": "</s>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "mask_token": {
24
+ "content": "<mask>",
25
+ "lstrip": true,
26
+ "normalized": false,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ },
30
+ "pad_token": {
31
+ "content": "<pad>",
32
+ "lstrip": false,
33
+ "normalized": false,
34
+ "rstrip": false,
35
+ "single_word": false
36
+ },
37
+ "sep_token": {
38
+ "content": "</s>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false
43
+ },
44
+ "unk_token": {
45
+ "content": "<unk>",
46
+ "lstrip": false,
47
+ "normalized": false,
48
+ "rstrip": false,
49
+ "single_word": false
50
+ }
51
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cad551d5600a84242d0973327029452a1e3672ba6313c2a3c3d69c4310e12719
3
+ size 17082987
tokenizer_config.json ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "<s>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "1": {
12
+ "content": "<pad>",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ },
19
+ "2": {
20
+ "content": "</s>",
21
+ "lstrip": false,
22
+ "normalized": false,
23
+ "rstrip": false,
24
+ "single_word": false,
25
+ "special": true
26
+ },
27
+ "3": {
28
+ "content": "<unk>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false,
33
+ "special": true
34
+ },
35
+ "250001": {
36
+ "content": "<mask>",
37
+ "lstrip": true,
38
+ "normalized": false,
39
+ "rstrip": false,
40
+ "single_word": false,
41
+ "special": true
42
+ }
43
+ },
44
+ "bos_token": "<s>",
45
+ "clean_up_tokenization_spaces": false,
46
+ "cls_token": "<s>",
47
+ "do_lower_case": true,
48
+ "eos_token": "</s>",
49
+ "extra_special_tokens": {},
50
+ "mask_token": "<mask>",
51
+ "max_length": 128,
52
+ "model_max_length": 128,
53
+ "pad_to_multiple_of": null,
54
+ "pad_token": "<pad>",
55
+ "pad_token_type_id": 0,
56
+ "padding_side": "right",
57
+ "sep_token": "</s>",
58
+ "stride": 0,
59
+ "strip_accents": null,
60
+ "tokenize_chinese_chars": true,
61
+ "tokenizer_class": "BertTokenizer",
62
+ "truncation_side": "right",
63
+ "truncation_strategy": "longest_first",
64
+ "unk_token": "<unk>"
65
+ }
unigram.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:da145b5e7700ae40f16691ec32a0b1fdc1ee3298db22a31ea55f57a966c4a65d
3
+ size 14763260