Serdar404 commited on
Commit
0bcd6fe
·
verified ·
1 Parent(s): 1f81119

Add RecGPT-10M model card

Browse files
Files changed (1) hide show
  1. README.md +56 -0
README.md CHANGED
@@ -1,3 +1,59 @@
1
  ---
 
 
2
  license: mit
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - en
4
  license: mit
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - babylm
9
+ - babylm-2026
10
+ - strict-small
11
+ - custom_code
12
  ---
13
+
14
+ # RecGPT-10M
15
+
16
+ RecGPT-10M is a 34.17M-parameter recursive causal language model trained for the BabyLM 2026 Strict-Small track. It was trained for 10 epochs on a custom 10M-word English corpus using a 32,768-token BPE vocabulary.
17
+
18
+ The model applies a shared Transformer block recursively for 16 iterations. Its hidden size is 768, embedding size is 192, and feed-forward intermediate size is 12,288. Training used Muon for the recursive block and AdamW for the embedding-related parameters, with a token batch size of 32,768 and sequence length 256.
19
+
20
+ ## Usage
21
+
22
+ This repository contains custom Transformers code, so loading requires `trust_remote_code=True`:
23
+
24
+ ```python
25
+ from transformers import AutoModelForCausalLM, AutoTokenizer
26
+
27
+ model_id = "Serdar404/RecGPT-10M"
28
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
29
+ model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
30
+ ```
31
+
32
+ The model is intended for scoring text as a causal language model. KV-cache generation is not currently implemented.
33
+
34
+ ## BabyLM 2026 evaluation
35
+
36
+ Final-checkpoint results before leaderboard collation:
37
+
38
+ | Evaluation | Score |
39
+ |---|---:|
40
+ | BLiMP | 73.11 |
41
+ | BLiMP Supplement | 61.73 |
42
+ | EWoK | 52.62 |
43
+ | Entity Tracking | 16.59 |
44
+ | COMPS | 55.43 |
45
+ | GlobalPIQA | 40.68 |
46
+ | (Super)GLUE | 66.64 |
47
+ | NLP Average | 52.40 |
48
+
49
+ Intermediate Strict-Small checkpoints are published as Hub revisions named `chck_1M` through `chck_100M` using the official BabyLM checkpoint schedule.
50
+
51
+ ## Resources
52
+
53
+ - Training code: https://github.com/serdardoesml/bblm26-recgpt
54
+ - Dataset construction: https://github.com/serdardoesml/bblm26-dataset
55
+ - Evaluation fork: https://github.com/serdardoesml/babylm-eval
56
+
57
+ ## Limitations
58
+
59
+ This is a small research model trained under the BabyLM data constraint. It is not intended for production deployment, factual question answering, or safety-critical use. Its outputs may contain inaccuracies or undesirable content inherited from its training data.