sciencemj's picture
Upload folder using huggingface_hub
938a6ab verified
|
Raw
History Blame Contribute Delete
3.2 kB
metadata
license: cdla-sharing-1.0
language:
  - en
library_name: pytorch
pipeline_tag: text-generation
datasets:
  - roneneldan/TinyStories
tags:
  - tinystories
  - small-language-model

tinyLLM 29M โ€” TinyStories

1๋‹จ๊ณ„ ์‚ฌ์ „ํ•™์Šต๋งŒ ๊ฑฐ์นœ ๊ฐ€์ค‘์น˜. ๋Œ€ํ™”๋Š” ๋ชป ํ•˜๊ณ  ์ด์•ผ๊ธฐ๋ฅผ ์ด์–ด์“ด๋‹ค.

ํŒŒ๋ผ๋ฏธํ„ฐ 29,577,728 ๊ฐœ. RTX 3060 Ti ํ•œ ๋Œ€์—์„œ ์‚ฌ์ „ํ•™์Šต 3.06 ์‹œ๊ฐ„ + SFT 2.4 ๋ถ„. ํ•™์Šต ์ฝ”๋“œ์™€ ์„ค๊ณ„ ๊ทผ๊ฑฐ: https://github.com/sciencemj/tinyLLM

val loss 1.3202 nats/token (perplexity 3.744, bits/char 0.4659)

์ด์šฉ ์กฐ๊ฑด

์ด์šฉ ์ œ์•ฝ์ด ์—†๋‹ค. TinyStories ๋Š” CDLA-Sharing-1.0 ์ด๊ณ  ยง3.5 ๊ฐ€ ๋ช…์‹œํ•œ๋‹ค โ€” "This Agreement imposes no obligations or restrictions on Your Use or Publication of Results." ยง1.11 ์—์„œ Results ๋Š” ๋ฐ์ดํ„ฐ์˜ Computational Use ๋กœ ์–ป์€ ์‚ฐ์ถœ๋ฌผ์ด๋ฉฐ, ์กฐ๊ฑด์€ ๋ฐ์ดํ„ฐ์˜ de minimis ๋ถ„๋Ÿ‰ ์ด์ƒ์„ ํฌํ•จํ•˜์ง€ ์•Š๋Š” ๊ฒƒ์ด๋‹ค. ์ด ๋ชจ๋ธ์€ train/val ๊ฒฉ์ฐจ๊ฐ€ 0.03 ์ด๋ผ ์ฝ”ํผ์Šค๋ฅผ ์™ธ์šฐ๊ณ  ์žˆ์ง€ ์•Š๋‹ค.

์“ฐ๋Š” ๋ฒ•

transformers ๋ฅผ ์“ฐ์ง€ ์•Š๋Š”๋‹ค. ์ด ์ €์žฅ์†Œ์˜ modeling_tinyllm.py ํ•˜๋‚˜๋ฉด ๋œ๋‹ค.

import torch
from tokenizers import Tokenizer
from modeling_tinyllm import TinyLM

model = TinyLM.from_pretrained(".")
tok = Tokenizer.from_file("tokenizer.json")

ids = torch.tensor([tok.encode("Once upon a time, there was a little girl named Lily.").ids])
out = model.generate(ids, 60, temperature=0.6, top_k=20)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))

์ด ๊ฐ€์ค‘์น˜๋Š” ๋Œ€ํ™”๋ฅผ ๋ชป ํ•œ๋‹ค. ์งˆ๋ฌธ์„ ์ฃผ๋ฉด ์ด์•ผ๊ธฐ์˜ ์ฒซ ๋ฌธ์žฅ์œผ๋กœ ๋ฐ›์•„ ๊ณ„์† ์จ ๋‚ด๋ ค๊ฐ„๋‹ค. ๋Œ€ํ™”๊ฐ€ ํ•„์š”ํ•˜๋ฉด tinyllm-29m-chat ์„ ์“ด๋‹ค.

๊ตฌ์กฐ

ids (B, 512)
  โ†’ nn.Embedding(8000, 512)  +  nn.Embedding(512, 512)
  โ†’ nn.TransformerEncoder(
        nn.TransformerEncoderLayer(512, nhead=8, dim_feedforward=2048,
                                   activation="gelu", norm_first=True,
                                   batch_first=True),
        num_layers=8, norm=nn.RMSNorm(512))
  โ†’ nn.Linear(512, 8000, bias=False)     # token embedding ๊ณผ tying

decoder-only ๋ฅผ TransformerEncoderLayer ๋กœ ๋งŒ๋“ ๋‹ค. TransformerDecoderLayer ๋Š” cross-attention ์šฉ memory ๋ฅผ ํ•„์ˆ˜๋กœ ์š”๊ตฌํ•ด์„œ ๋งž์ง€ ์•Š๋Š”๋‹ค.

ํ† ํฌ๋‚˜์ด์ €๋Š” TinyStories ์™€ DailyDialog ํ•ฉ์ง‘ํ•ฉ์—์„œ ํ•™์Šตํ•œ ์ž์ฒด 8k byte-level BPE ๋‹ค. ๊ฐ™์ด ๋ฐ›์€ tokenizer.json ์„ ๋ฐ˜๋“œ์‹œ ์จ์•ผ ํ•œ๋‹ค. ๋‹ค๋ฅธ ํ† ํฌ๋‚˜์ด์ €๋กœ๋Š” ๋™์ž‘ํ•˜์ง€ ์•Š๋Š”๋‹ค.

ํ•œ๊ณ„

๋œ๋‹ค โ€” ๋ฌธ๋ฒ•, ๊ตฌ๋‘์ , ๋”ฐ์˜ดํ‘œ ๋Œ€ํ™” ํ˜•์‹, ๋ฌธ๋‹จ ๋‚˜๋ˆ„๊ธฐ, ์ธ๋ฌผ ์ด๋ฆ„ ์œ ์ง€, ์ธ๊ณผ ์—ฐ๊ฒฐ.

์•ˆ ๋œ๋‹ค โ€” ํ„ด ๊ฐ„ ๊ธฐ์–ต, ์งˆ๋ฌธ์— ๋Œ€ํ•œ ์ง์ ‘ ๋‹ต๋ณ€, ์‚ฌ์‹ค์„ฑ, ๋ฌธ์žฅ ์•ˆ ๋ฐ˜๋ณต, ๋…ผ๋ฆฌ ์ผ๊ด€์„ฑ. ์˜์–ด๋งŒ ์•ˆ๋‹ค. ์‚ฌ์‹ค ์ •๋ณด๋ฅผ ์–ป๋Š” ์šฉ๋„๋กœ ์“ฐ๋ฉด ์•ˆ ๋œ๋‹ค.

์ž์„ธํ•œ ๊ฒƒ์€ MODEL_CARD.md.

์ธ์šฉ

@article{eldan2023tinystories,
  title={TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
  author={Eldan, Ronen and Li, Yuanzhi},
  journal={arXiv preprint arXiv:2305.07759},
  year={2023}
}