--- license: mit tags: - transformer - from-scratch - pytorch - text-generation datasets: - roneneldan/TinyStories --- # Mini Transformer (From Scratch) A decoder-only transformer built entirely from first principles in PyTorch, no pre-built attention layers, no HuggingFace Transformers library internals. Every component (multi-head self-attention, positional encoding, feedforward blocks, residual connections) is implemented from scratch and trained on the TinyStories dataset. This is part of a larger project studying transformer internals, followed by implementing DeepSeek's MLA (Multi-head Latent Attention) and mHC (Manifold-Constrained Hyper-Connections) as architectural upgrades to this same baseline. ## Architecture - Decoder-only transformer, GPT-style - Embedding dim: 256 - Attention heads: 8 - Layers: 6 - Feedforward hidden dim: 1024 - Context length: 128 tokens - Tokenizer: GPT-2 (tiktoken), 50,257 vocab - Total parameters: 336,488 ## Training - Dataset: TinyStories (roneneldan/TinyStories), ~474M training tokens - Steps: 10,000 - Batch size: 32 - Optimizer: AdamW, lr=3e-4 - Hardware: single RTX 3060 (6GB VRAM) - Final train loss: None - Final val loss: None ## Usage Load `model.py` for the architecture classes, then: ```python import torch from model import MiniGPT model = MiniGPT(vocab_size=50257, embed_dim=256, num_heads=8, num_layers=6, hidden_dim=1024, max_seq_len=128) ckpt = torch.load("checkpoint.pt", map_location="cpu") model.load_state_dict(ckpt["model"]) ``` ## Sample Output > Once upon a time, there was a little girl named Lily who loved to play > outside. She had a lot of fun playing in the mud and splashing around...