Update README.md
#1
by Datdanboi25 - opened
README.md
CHANGED
|
@@ -42,6 +42,7 @@ We tested two different architecture configurations on 500M tokens of FineWeb to
|
|
| 42 |
- Pretraining Tp: `1`
|
| 43 |
- Use Cache: `false`
|
| 44 |
- **Total Parameters**: `3,889,920`
|
|
|
|
| 45 |
|
| 46 |
### Config B: Shallow-Wide (code name: `width_311`)
|
| 47 |
|
|
@@ -65,6 +66,7 @@ We tested two different architecture configurations on 500M tokens of FineWeb to
|
|
| 65 |
- Pretraining Tp: `1`
|
| 66 |
- Use Cache: `false`
|
| 67 |
- **Total Parameters**: `3,816,576`
|
|
|
|
| 68 |
|
| 69 |
## Training Setup
|
| 70 |
|
|
|
|
| 42 |
- Pretraining Tp: `1`
|
| 43 |
- Use Cache: `false`
|
| 44 |
- **Total Parameters**: `3,889,920`
|
| 45 |
+
- **Hidden Size/Layers**: `6.10`
|
| 46 |
|
| 47 |
### Config B: Shallow-Wide (code name: `width_311`)
|
| 48 |
|
|
|
|
| 66 |
- Pretraining Tp: `1`
|
| 67 |
- Use Cache: `false`
|
| 68 |
- **Total Parameters**: `3,816,576`
|
| 69 |
+
- **Hidden Size/Layers**: `21.33`
|
| 70 |
|
| 71 |
## Training Setup
|
| 72 |
|