prathamkode commited on
Commit
1e1c225
·
verified ·
1 Parent(s): 5c5b094

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +66 -65
README.md CHANGED
@@ -10,70 +10,71 @@ tags:
10
  - smol
11
  datasets:
12
  - HuggingFaceTB/smol-smoltalk
 
13
  ---
14
-
15
- # particle-1.0
16
-
17
- ~100M-parameter Llama-style chat model trained **from scratch** (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.
18
-
19
- Weights are MIT. Training data still needs attribution (below).
20
-
21
- ## Usage
22
-
23
- ```python
24
- from transformers import AutoModelForCausalLM, AutoTokenizer
25
-
26
- repo = "prathamkode/particle-1.0"
27
- tok = AutoTokenizer.from_pretrained(repo)
28
- model = AutoModelForCausalLM.from_pretrained(repo)
29
-
30
- messages = [{"role": "user", "content": "hello"}]
31
- prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
32
- ids = tok(prompt, return_tensors="pt")
33
- out = model.generate(**ids, max_new_tokens=64, temperature=0.7)
34
- print(tok.decode(out[0], skip_special_tokens=False))
35
- ```
36
-
37
- Chat format:
38
-
39
- ```
40
- <|user|>
41
- hello
42
- <|assistant|>
43
- ```
44
-
45
- ## Model details
46
-
47
- | | |
48
- |---|---|
49
- | Architecture | Llama-style decoder (RoPE, SwiGLU, RMSNorm, tied embeddings) |
50
- | Parameters | ~100M (12 layers, 768 hidden, 12 heads) |
51
- | Context | 2048 tokens |
52
- | Tokenizer | Custom 32k byte-level BPE (not Llama / GPT-2 vocab) |
53
- | Init | Random `N(0, 0.02)` — trained from scratch |
54
- | Precision | BF16 training; Hub weights `bfloat16` |
55
-
56
- ## Training
57
-
58
- 1. **Tokenizer** trained from scratch on a FineWeb-Edu sample (~2GB text).
59
- 2. **Pretrain** next-token prediction on [`HuggingFaceFW/fineweb_edu_100BT-shuffled`](https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled), first ~2B tokens.
60
- 3. **SFT** on [`HuggingFaceTB/smol-smoltalk`](https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk) (first user/assistant turn + a few greeting seeds).
61
-
62
- SFT used that dataset as **text only**. No teacher model weights were copied.
63
-
64
- ## Intended use
65
-
66
- Research / demo small chat model. Expect short replies, mistakes, and weak reasoning.
67
-
68
- ## Limitations
69
-
70
- - Very small capacity
71
- - May hallucinate
72
- - English-centric FineWeb-Edu subset
73
- - No RLHF / preference tuning
74
-
75
- ## License
76
-
77
- - **These weights:** [MIT](LICENSE)
78
- - **FineWeb-Edu:** ODC-By (attribute)
79
  - **smol-smoltalk:** follow the dataset card
 
10
  - smol
11
  datasets:
12
  - HuggingFaceTB/smol-smoltalk
13
+ - HuggingFaceFW/fineweb_edu_100BT-shuffled
14
  ---
15
+
16
+ # particle-1.0
17
+
18
+ ~100M-parameter Llama-style chat model trained **from scratch** (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.
19
+
20
+ Weights are MIT. Training data still needs attribution (below).
21
+
22
+ ## Usage
23
+
24
+ ```python
25
+ from transformers import AutoModelForCausalLM, AutoTokenizer
26
+
27
+ repo = "prathamkode/particle-1.0"
28
+ tok = AutoTokenizer.from_pretrained(repo)
29
+ model = AutoModelForCausalLM.from_pretrained(repo)
30
+
31
+ messages = [{"role": "user", "content": "hello"}]
32
+ prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
33
+ ids = tok(prompt, return_tensors="pt")
34
+ out = model.generate(**ids, max_new_tokens=64, temperature=0.7)
35
+ print(tok.decode(out[0], skip_special_tokens=False))
36
+ ```
37
+
38
+ Chat format:
39
+
40
+ ```
41
+ <|user|>
42
+ hello
43
+ <|assistant|>
44
+ ```
45
+
46
+ ## Model details
47
+
48
+ | | |
49
+ |---|---|
50
+ | Architecture | Llama-style decoder (RoPE, SwiGLU, RMSNorm, tied embeddings) |
51
+ | Parameters | ~100M (12 layers, 768 hidden, 12 heads) |
52
+ | Context | 2048 tokens |
53
+ | Tokenizer | Custom 32k byte-level BPE (not Llama / GPT-2 vocab) |
54
+ | Init | Random `N(0, 0.02)` — trained from scratch |
55
+ | Precision | BF16 training; Hub weights `bfloat16` |
56
+
57
+ ## Training
58
+
59
+ 1. **Tokenizer** trained from scratch on a FineWeb-Edu sample (~2GB text).
60
+ 2. **Pretrain** next-token prediction on [`HuggingFaceFW/fineweb_edu_100BT-shuffled`](https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled), first ~2B tokens.
61
+ 3. **SFT** on [`HuggingFaceTB/smol-smoltalk`](https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk) (first user/assistant turn + a few greeting seeds).
62
+
63
+ SFT used that dataset as **text only**. No teacher model weights were copied.
64
+
65
+ ## Intended use
66
+
67
+ Research / demo small chat model. Expect short replies, mistakes, and weak reasoning.
68
+
69
+ ## Limitations
70
+
71
+ - Very small capacity
72
+ - May hallucinate
73
+ - English-centric FineWeb-Edu subset
74
+ - No RLHF / preference tuning
75
+
76
+ ## License
77
+
78
+ - **These weights:** [MIT](LICENSE)
79
+ - **FineWeb-Edu:** ODC-By (attribute)
80
  - **smol-smoltalk:** follow the dataset card