deepset
/

tinybert-6l-768d-squad2

Question Answering

Eval Results (legacy)

Model card Files Files and versions

MichelBartelsDeepset commited on Jan 14, 2022

Commit

6c62aa0

·

1 Parent(s): 8084f28

Update README.md

Files changed (1) hide show

README.md +3 -4

README.md CHANGED Viewed

@@ -11,13 +11,13 @@ tags:
 ## Overview
 **Language model:** deepset/tinybert-6L-768D-squad2
 **Language:** English
-**Training data:** SQuAD 2.0 training set x 20 augmented + SQuAD 2.0 training set
 **Eval data:** SQuAD 2.0 dev set
 **Infrastructure**: 1x V100 GPU
 **Published**: Dec 8th, 2021
 ## Details
-- haystack's intermediate layer and prediction layer distillation features were used for training (based on [TinyBERT](https://arxiv.org/pdf/1909.10351.pdf)). deepset/bert-base-uncased-squad2 was used as the teacher model.
 ## Hyperparameters
 ### Intermediate layer distillation
@@ -29,7 +29,6 @@ learning_rate = 5e-5
 lr_schedule = LinearWarmup
 embeds_dropout_prob = 0.1
 temperature = 1
-distillation_loss_weight = 0.75
 ```
 ### Prediction layer distillation
 ```
@@ -40,7 +39,7 @@ learning_rate = 3e-5
 lr_schedule = LinearWarmup
 embeds_dropout_prob = 0.1
 temperature = 1
-distillation_loss_weight = 0.75
 ```
 ## Performance
 ```

 ## Overview
 **Language model:** deepset/tinybert-6L-768D-squad2
 **Language:** English
+**Training data:** SQuAD 2.0 training set x 20 augmented + SQuAD 2.0 training set without augmentation
 **Eval data:** SQuAD 2.0 dev set
 **Infrastructure**: 1x V100 GPU
 **Published**: Dec 8th, 2021
 ## Details
+- haystack's intermediate layer and prediction layer distillation features were used for training (based on [TinyBERT](https://arxiv.org/pdf/1909.10351.pdf)). deepset/bert-base-uncased-squad2 was used as the teacher model and huawei-noah/TinyBERT_General_6L_768D was used as the student model.
 ## Hyperparameters
 ### Intermediate layer distillation
 lr_schedule = LinearWarmup
 embeds_dropout_prob = 0.1
 temperature = 1
 ```
 ### Prediction layer distillation
 ```
 lr_schedule = LinearWarmup
 embeds_dropout_prob = 0.1
 temperature = 1
+distillation_loss_weight = 1.0
 ```
 ## Performance
 ```