--- language: en tags: - text-generation - gpt2 - from-scratch - cpu-training - instruction-tuning datasets: - tatsu-lab/alpaca inference: true --- # CPU Chat LLM A GPT-2 family language model trained **entirely from scratch on CPU** using the Alpaca instruction dataset. This project demonstrates training a choppable language model with limited compute resources. ## Training History | Run | Model | Params | Training Data | Time | Steps | Epochs | Min Loss | |---|---|---|---|---|---|---|---| | 1 | 51M Base | 51.1M | Full Alpaca (52k) | 28 min | 848 | ~0.02 | 2.85 | | 2 | 51M v2 | 51.1M | 500 short examples | 10 min | ~350 | 7 | 2.07 | | 3 | 16M Chat | 16.1M | 500 short examples | 10 min | 460 | 16 | 1.98 | | 4 | 51M v3 | 51.1M | 10k filtered Alpaca | 10 min+ | in-progress | - | - | All models were initialized with random weights and trained from scratch — no pre-training or transfer learning. ## Available Models on Hugging Face | Model ID | Params | Size | Description | |---|---|---|---| | `USAGAMES365/cpu-chat-llm/models/v2-51m` | 51.1M | 204 MB | Retrained 500 short examples, 7 epochs | | `USAGAMES365/cpu-chat-llm/models/v3-51m` | 51.1M | 204 MB | Continued training on 10k examples (in progress) | | `USAGAMES365/cpu-chat-llm/models/gemma-270m` | 268M | 1.0 GB | Gemma-3-270m-it for comparison (pre-trained, not ours) | ## Usage ```python from transformers import GPT2LMHeadModel, AutoTokenizer # Load our model model = GPT2LMHeadModel.from_pretrained("USAGAMES365/cpu-chat-llm/models/v2-51m") tokenizer = AutoTokenizer.from_pretrained("USAGAMES365/cpu-chat-llm/models/v2-51m") # Format prompt correctly prompt = "<|User|> What is 2+2?<|Assistant|>" inputs = tokenizer.encode(prompt, return_tensors="pt") # Generate outputs = model.generate(inputs, max_new_tokens=80, temperature=0.7, do_sample=True) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## Training Details - **Architecture:** GPT-2 decoder-only transformer - **Config (51M):** 8 layers, 8 heads, 512 hidden dim, 2048 FF dim, 192-256 context - **Config (16M):** 6 layers, 4 heads, 256 hidden dim, 512 FF dim, 192 context - **Tokenizer:** GPT-2 tokenizer (50,257 vocab) - **Optimizer:** AdamW (lr=3e-4 to 1e-3, weight_decay=0.01) - **Batch size:** 4-16 - **Hardware:** CPU only (Replit environment) - **Software:** PyTorch, Transformers, Datasets ## Run Structure The repo organizes checkpoints from each training phase: - `previous-run/` - Initial training runs with step-level checkpoints - `current-run/` - Continued training with saved checkpoints - `models/` - Final exportable models ## Limitations - Small model size (16M-51M params vs. billions in production models) - Very limited training time (minutes vs. thousands of GPU-hours) - CPU-only training constrains both speed and model capacity - Outputs may be incoherent or repetitive - Not suitable for production deployment