w-ahmad commited on
Commit
e210adb
·
verified ·
1 Parent(s): ffdfc0d

End of training

Browse files
Files changed (6) hide show
  1. .gitattributes +1 -0
  2. README.md +23 -23
  3. config.json +1 -1
  4. model.safetensors +1 -1
  5. training_args.bin +1 -1
  6. training_log.jsonl +0 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ training_log.jsonl filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -14,7 +14,7 @@ should probably proofread and complete it, then remove this comment. -->
14
 
15
  This model is a fine-tuned version of [](https://huggingface.co/) on an unknown dataset.
16
  It achieves the following results on the evaluation set:
17
- - Loss: 0.0753
18
 
19
  ## Model description
20
 
@@ -33,7 +33,7 @@ More information needed
33
  ### Training hyperparameters
34
 
35
  The following hyperparameters were used during training:
36
- - learning_rate: 0.0005
37
  - train_batch_size: 64
38
  - eval_batch_size: 64
39
  - seed: 42
@@ -45,27 +45,27 @@ The following hyperparameters were used during training:
45
 
46
  | Training Loss | Epoch | Step | Validation Loss |
47
  |:-------------:|:------:|:----:|:---------------:|
48
- | 1.6407 | 0.0135 | 200 | 1.4343 |
49
- | 0.2608 | 0.0270 | 400 | 0.2564 |
50
- | 0.1712 | 0.0404 | 600 | 0.1649 |
51
- | 0.1275 | 0.0539 | 800 | 0.1277 |
52
- | 0.1260 | 0.0674 | 1000 | 0.1178 |
53
- | 0.1029 | 0.0809 | 1200 | 0.1025 |
54
- | 0.0961 | 0.0944 | 1400 | 0.0965 |
55
- | 0.1125 | 0.1079 | 1600 | 0.1061 |
56
- | 0.0891 | 0.1213 | 1800 | 0.0886 |
57
- | 0.0847 | 0.1348 | 2000 | 0.0851 |
58
- | 0.0824 | 0.1483 | 2200 | 0.0831 |
59
- | 0.0841 | 0.1618 | 2400 | 0.0828 |
60
- | 0.0797 | 0.1753 | 2600 | 0.0801 |
61
- | 0.0787 | 0.1887 | 2800 | 0.0791 |
62
- | 0.0772 | 0.2022 | 3000 | 0.0781 |
63
- | 0.0770 | 0.2157 | 3200 | 0.0775 |
64
- | 0.0763 | 0.2292 | 3400 | 0.0769 |
65
- | 0.0761 | 0.2427 | 3600 | 0.0764 |
66
- | 0.0755 | 0.2562 | 3800 | 0.0759 |
67
- | 0.0755 | 0.2696 | 4000 | 0.0755 |
68
- | 0.0750 | 0.2761 | 4096 | 0.0753 |
69
 
70
 
71
  ### Framework versions
 
14
 
15
  This model is a fine-tuned version of [](https://huggingface.co/) on an unknown dataset.
16
  It achieves the following results on the evaluation set:
17
+ - Loss: 0.1862
18
 
19
  ## Model description
20
 
 
33
  ### Training hyperparameters
34
 
35
  The following hyperparameters were used during training:
36
+ - learning_rate: 0.0003
37
  - train_batch_size: 64
38
  - eval_batch_size: 64
39
  - seed: 42
 
45
 
46
  | Training Loss | Epoch | Step | Validation Loss |
47
  |:-------------:|:------:|:----:|:---------------:|
48
+ | 3.3162 | 0.0135 | 200 | 3.1383 |
49
+ | 0.8013 | 0.0270 | 400 | 0.7547 |
50
+ | 0.4088 | 0.0404 | 600 | 0.4044 |
51
+ | 0.3265 | 0.0539 | 800 | 0.3272 |
52
+ | 0.2864 | 0.0674 | 1000 | 0.2859 |
53
+ | 0.2629 | 0.0809 | 1200 | 0.2619 |
54
+ | 0.2403 | 0.0944 | 1400 | 0.2424 |
55
+ | 0.2472 | 0.1079 | 1600 | 0.2468 |
56
+ | 0.2213 | 0.1213 | 1800 | 0.2201 |
57
+ | 0.2111 | 0.1348 | 2000 | 0.2126 |
58
+ | 0.2040 | 0.1483 | 2200 | 0.2070 |
59
+ | 0.2141 | 0.1618 | 2400 | 0.2089 |
60
+ | 0.1972 | 0.1753 | 2600 | 0.1989 |
61
+ | 0.1940 | 0.1887 | 2800 | 0.1959 |
62
+ | 0.1902 | 0.2022 | 3000 | 0.1934 |
63
+ | 0.1898 | 0.2157 | 3200 | 0.1916 |
64
+ | 0.1878 | 0.2292 | 3400 | 0.1901 |
65
+ | 0.1874 | 0.2427 | 3600 | 0.1888 |
66
+ | 0.1860 | 0.2562 | 3800 | 0.1876 |
67
+ | 0.1863 | 0.2696 | 4000 | 0.1866 |
68
+ | 0.1849 | 0.2761 | 4096 | 0.1862 |
69
 
70
 
71
  ### Framework versions
config.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "activation": "gelu",
3
  "architectures": [
4
  "TinyLlamaForCausalLM"
5
  ],
 
1
  {
2
+ "activation": "silu",
3
  "architectures": [
4
  "TinyLlamaForCausalLM"
5
  ],
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e41744cf69ceba537aa4065b08b4c0d17890e3277555fac409a90551bce0ab80
3
  size 4011496
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:59507337f6067a80b43664e709b7c20c03e78f5deeb583a8dc3f447480ecc287
3
  size 4011496
training_args.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:51de4cb4b05849f4f3158b0962940d33e0f83f023656fcb707723ff7fd2c765e
3
  size 4856
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:90c2469c0f3298da74de93055356d686cfe87200dcf88f8003df31a102ba07f2
3
  size 4856
training_log.jsonl CHANGED
The diff for this file is too large to render. See raw diff