| ---
|
| license: mit
|
| tags:
|
| - transformer
|
| - from-scratch
|
| - pytorch
|
| - text-generation
|
| datasets:
|
| - roneneldan/TinyStories
|
| ---
|
|
|
| # Mini Transformer (From Scratch)
|
|
|
| A decoder-only transformer built entirely from first principles in PyTorch,
|
| no pre-built attention layers, no HuggingFace Transformers library internals.
|
| Every component (multi-head self-attention, positional encoding, feedforward
|
| blocks, residual connections) is implemented from scratch and trained on the
|
| TinyStories dataset.
|
|
|
| This is part of a larger project studying transformer internals, followed by
|
| implementing DeepSeek's MLA (Multi-head Latent Attention) and mHC
|
| (Manifold-Constrained Hyper-Connections) as architectural upgrades to this
|
| same baseline.
|
|
|
| ## Architecture
|
|
|
| - Decoder-only transformer, GPT-style
|
| - Embedding dim: 256
|
| - Attention heads: 8
|
| - Layers: 6
|
| - Feedforward hidden dim: 1024
|
| - Context length: 128 tokens
|
| - Tokenizer: GPT-2 (tiktoken), 50,257 vocab
|
| - Total parameters: 336,488
|
|
|
| ## Training
|
|
|
| - Dataset: TinyStories (roneneldan/TinyStories), ~474M training tokens
|
| - Steps: 10,000
|
| - Batch size: 32
|
| - Optimizer: AdamW, lr=3e-4
|
| - Hardware: single RTX 3060 (6GB VRAM)
|
| - Final train loss: None
|
| - Final val loss: None
|
|
|
| ## Usage
|
|
|
| Load `model.py` for the architecture classes, then:
|
|
|
| ```python
|
| import torch
|
| from model import MiniGPT
|
|
|
| model = MiniGPT(vocab_size=50257, embed_dim=256, num_heads=8,
|
| num_layers=6, hidden_dim=1024, max_seq_len=128)
|
| ckpt = torch.load("checkpoint.pt", map_location="cpu")
|
| model.load_state_dict(ckpt["model"])
|
| ```
|
|
|
| ## Sample Output
|
|
|
| > Once upon a time, there was a little girl named Lily who loved to play
|
| > outside. She had a lot of fun playing in the mud and splashing around...
|
|
|