license: cdla-sharing-1.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
datasets:
- roneneldan/TinyStories
tags:
- tinystories
- small-language-model
tinyLLM 29M โ TinyStories
1๋จ๊ณ ์ฌ์ ํ์ต๋ง ๊ฑฐ์น ๊ฐ์ค์น. ๋ํ๋ ๋ชป ํ๊ณ ์ด์ผ๊ธฐ๋ฅผ ์ด์ด์ด๋ค.
ํ๋ผ๋ฏธํฐ 29,577,728 ๊ฐ. RTX 3060 Ti ํ ๋์์ ์ฌ์ ํ์ต 3.06 ์๊ฐ + SFT 2.4 ๋ถ. ํ์ต ์ฝ๋์ ์ค๊ณ ๊ทผ๊ฑฐ: https://github.com/sciencemj/tinyLLM
val loss 1.3202 nats/token (perplexity 3.744, bits/char 0.4659)
์ด์ฉ ์กฐ๊ฑด
์ด์ฉ ์ ์ฝ์ด ์๋ค. TinyStories ๋ CDLA-Sharing-1.0 ์ด๊ณ ยง3.5 ๊ฐ ๋ช ์ํ๋ค โ "This Agreement imposes no obligations or restrictions on Your Use or Publication of Results." ยง1.11 ์์ Results ๋ ๋ฐ์ดํฐ์ Computational Use ๋ก ์ป์ ์ฐ์ถ๋ฌผ์ด๋ฉฐ, ์กฐ๊ฑด์ ๋ฐ์ดํฐ์ de minimis ๋ถ๋ ์ด์์ ํฌํจํ์ง ์๋ ๊ฒ์ด๋ค. ์ด ๋ชจ๋ธ์ train/val ๊ฒฉ์ฐจ๊ฐ 0.03 ์ด๋ผ ์ฝํผ์ค๋ฅผ ์ธ์ฐ๊ณ ์์ง ์๋ค.
์ฐ๋ ๋ฒ
transformers ๋ฅผ ์ฐ์ง ์๋๋ค. ์ด ์ ์ฅ์์ modeling_tinyllm.py ํ๋๋ฉด ๋๋ค.
import torch
from tokenizers import Tokenizer
from modeling_tinyllm import TinyLM
model = TinyLM.from_pretrained(".")
tok = Tokenizer.from_file("tokenizer.json")
ids = torch.tensor([tok.encode("Once upon a time, there was a little girl named Lily.").ids])
out = model.generate(ids, 60, temperature=0.6, top_k=20)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
์ด ๊ฐ์ค์น๋ ๋ํ๋ฅผ ๋ชป ํ๋ค. ์ง๋ฌธ์ ์ฃผ๋ฉด ์ด์ผ๊ธฐ์ ์ฒซ ๋ฌธ์ฅ์ผ๋ก ๋ฐ์ ๊ณ์ ์จ ๋ด๋ ค๊ฐ๋ค. ๋ํ๊ฐ ํ์ํ๋ฉด tinyllm-29m-chat ์ ์ด๋ค.
๊ตฌ์กฐ
ids (B, 512)
โ nn.Embedding(8000, 512) + nn.Embedding(512, 512)
โ nn.TransformerEncoder(
nn.TransformerEncoderLayer(512, nhead=8, dim_feedforward=2048,
activation="gelu", norm_first=True,
batch_first=True),
num_layers=8, norm=nn.RMSNorm(512))
โ nn.Linear(512, 8000, bias=False) # token embedding ๊ณผ tying
decoder-only ๋ฅผ TransformerEncoderLayer ๋ก ๋ง๋ ๋ค. TransformerDecoderLayer ๋
cross-attention ์ฉ memory ๋ฅผ ํ์๋ก ์๊ตฌํด์ ๋ง์ง ์๋๋ค.
ํ ํฌ๋์ด์ ๋ TinyStories ์ DailyDialog ํฉ์งํฉ์์ ํ์ตํ ์์ฒด 8k byte-level BPE ๋ค.
๊ฐ์ด ๋ฐ์ tokenizer.json ์ ๋ฐ๋์ ์จ์ผ ํ๋ค. ๋ค๋ฅธ ํ ํฌ๋์ด์ ๋ก๋ ๋์ํ์ง ์๋๋ค.
ํ๊ณ
๋๋ค โ ๋ฌธ๋ฒ, ๊ตฌ๋์ , ๋ฐ์ดํ ๋ํ ํ์, ๋ฌธ๋จ ๋๋๊ธฐ, ์ธ๋ฌผ ์ด๋ฆ ์ ์ง, ์ธ๊ณผ ์ฐ๊ฒฐ.
์ ๋๋ค โ ํด ๊ฐ ๊ธฐ์ต, ์ง๋ฌธ์ ๋ํ ์ง์ ๋ต๋ณ, ์ฌ์ค์ฑ, ๋ฌธ์ฅ ์ ๋ฐ๋ณต, ๋ ผ๋ฆฌ ์ผ๊ด์ฑ. ์์ด๋ง ์๋ค. ์ฌ์ค ์ ๋ณด๋ฅผ ์ป๋ ์ฉ๋๋ก ์ฐ๋ฉด ์ ๋๋ค.
์์ธํ ๊ฒ์ MODEL_CARD.md.
์ธ์ฉ
@article{eldan2023tinystories,
title={TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
author={Eldan, Ronen and Li, Yuanzhi},
journal={arXiv preprint arXiv:2305.07759},
year={2023}
}