Safetensors
English
llama
loubnabnl HF Staff commited on
Commit
00337b0
·
verified ·
1 Parent(s): 9f54686

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +14 -1
README.md CHANGED
@@ -62,4 +62,17 @@ print(tokenizer.decode(outputs[0]))
62
  We used the SmolLM2 setup to evaluate all our ablation models with `lighteval`. You can find the details here: https://github.com/huggingface/smollm/tree/main/evaluation#smollm2-base-models
63
 
64
  ## Limitations
65
- This model was predominantly trained on English math data, potentially limiting its performance in other languages. Furthermore, the model's behavior is influenced by the quality and diversity of its training data, which may include biases and harmful content.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
  We used the SmolLM2 setup to evaluate all our ablation models with `lighteval`. You can find the details here: https://github.com/huggingface/smollm/tree/main/evaluation#smollm2-base-models
63
 
64
  ## Limitations
65
+ This model was predominantly trained on English math data, potentially limiting its performance in other languages. Furthermore, the model's behavior is influenced by the quality and diversity of its training data, which may include biases and harmful content.
66
+
67
+ ## Citation
68
+ ```bash
69
+ @misc{allal2025smollm2smolgoesbig,
70
+ title={SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model},
71
+ author={Loubna Ben Allal and Anton Lozhkov and Elie Bakouch and Gabriel Martín Blázquez and Guilherme Penedo and Lewis Tunstall and Andrés Marafioti and Hynek Kydlíček and Agustín Piqueres Lajarín and Vaibhav Srivastav and Joshua Lochner and Caleb Fahlgren and Xuan-Son Nguyen and Clémentine Fourrier and Ben Burtenshaw and Hugo Larcher and Haojun Zhao and Cyril Zakka and Mathieu Morlon and Colin Raffel and Leandro von Werra and Thomas Wolf},
72
+ year={2025},
73
+ eprint={2502.02737},
74
+ archivePrefix={arXiv},
75
+ primaryClass={cs.CL},
76
+ url={https://arxiv.org/abs/2502.02737},
77
+ }
78
+ ```