| datasets: | |
| - roneneldan/TinyStories | |
| metrics: | |
| - babylm | |
| Basemodel: roBERTa | |
| Configs: | |
| Vocab size: 10,000 | |
| Hidden size: 512 | |
| Max position embeddings: 512 | |
| Number of layers: 2 | |
| Number of heads: 4 | |
| Window size: 256 | |
| Intermediate-size: 1024 | |
| Results: | |
| - Task: glue | |
| Score: 57.54 | |
| Confidence Interval: [57.15, 58.1] | |
| - Task: blimp | |
| Score: 59.16 | |
| Confidence Interval: [58.75, 59.53] | |