--- license: mit language: - en pipeline_tag: text-generation tags: - causal-lm - pytorch - small-language-model - autoresearch - from-scratch datasets: - tinystories --- # AutoResearch-tinystories-depth8 ![AutoResearch Cover](https://raw.githubusercontent.com/nullbotai-droid/turbo-diffusion-local/refs/heads/main/cover2.png) **AutoResearch-tinystories-depth8** is a **285.2M parameter** decoder-only Transformer trained **from scratch** on **TinyStories (karpathy/tinystories-gpt4-clean)**. This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows. --- # Overview This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.5 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.509277 (perplexity: 1.4233) on the held-out validation set. --- # References ## Papers - NanoGPT / NanoChat architecture patterns ## Datasets - [TinyStories (karpathy/tinystories-gpt4-clean)](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean) ## Related Projects - [karpathy/autoresearch](https://github.com/karpathy/autoresearch) ## WANDB Run - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1) --- # Highlights - Trained **from scratch** - **285.2M parameters** - Trained on **39.8M tokens** (76 steps) - 8-layer decoder-only Transformer with sliding window attention - RoPE positional encoding, RMSNorm, ReLU² activation - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings) - **Mixture of Experts** (8 routed + 1 shared, top-2 routing) - Hugging Face Transformers compatible --- # Model Architecture | Property | Value | |-----------|------:| | Architecture | Decoder-only Transformer | | Parameters | **285,245,968** (285.2M) | | Layers | 8 | | Hidden Size | 512 | | Attention Heads | 4 | | KV Heads | 4 | | Head Dimension | 128 | | Feed Forward Size | 2048 (MoE: 8 experts, 1 shared, top-2) | | Context Length | 2048 | | Vocabulary Size | 16,384 | | Positional Encoding | RoPE | | Activation | ReLU² | | Normalization | RMSNorm | | Window Pattern | SSSL | | Weight Tying | No | --- # Training This model was trained **from scratch** for **0.5 hours** (1802s) of wall-clock training time. ## Training Configuration | Setting | Value | |---------|------:| | Optimizer | MuonAdamW (Muon + AdamW) | | Precision | torch.bfloat16 | | Learning Rate | 0.04 (matrix) / 0.6 (embedding) | | Weight Decay | 0.2 | | Batch Size | 4 × 2048 = 8,192 tokens/step | | Gradient Accumulation | 64 steps | | Total Batch Size | 524,288 tokens | | Context Length | 2048 | | Vocabulary | 16,384 tokens (BPE) | | LR Scheduler | Linear warmdown (50%) | | Activation Checkpointing | Enabled | ## Hardware - GPU: NVIDIA GeForce RTX 4060 Ti - VRAM: 16.0 GB - Peak VRAM Used: 6.3 GB - MFU: 33.08% - Framework: PyTorch 2.9.1+cu128 --- # Dataset - **Name:** TinyStories (karpathy/tinystories-gpt4-clean) - **Language:** English ### Preprocessing Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset. --- # Intended Use This model is intended for: - Educational purposes and research - Text generation experiments - Studying small language model training dynamics Not recommended for: - Production use or safety-critical applications - Tasks requiring factual accuracy --- # Evaluation ## Results | Metric | Score | |---------|------:| | Validation BPB | 0.509277 | | Perplexity | 1.4233 | | Peak VRAM | 6.3 GB | | MFU | 33.08% | --- # Example Generations ## Example 1 ### Prompt ```text Once upon a time, ``` ### Generation ```text Once upon a time, 3 year old boy, Jack, went outside to play. He saw a black and white black dog running around. He pointed and said, "Hi! I'm Jack, this is my dog. He is very black!" Jack smiled and said, ``` ## Example 2 ### Prompt ```text A lonely dragon ``` ### Generation ```text A lonely dragon 3 year old child exploring the new places. She had no one to come along, and no one else was ever seen in the world.<|bos|>Once upon a time, there was a little boy named Tim. Tim had a long rod that he loved ``` ## Example 3 ### Prompt ```text The opposite of boy is ``` ### Generation ```text The opposite of boy is 3 years old. The boy said "I can help you, I can do anything you want. That's very kind of you." The boy smiled and said "Thank you, boy. I'm so happy to help. Thank you, you're ``` ## Example 4 ### Prompt ```text The opposite of queen is ``` ### Generation ```text The opposite of queen is 3 and she is very happy. She hugs queen and kisses her cheek. queen and queen queen were very happy.<|bos|>Once upon a time, there was a little dog named Tim. Tim loved to play in the park. One day, while playing ``` ## Example 5 ### Prompt ```text My name is ``` ### Generation ```text My name is 4rayla and I am 3 years 3-year-old. He was very happy with his new friend. He was always excited to learn about how to act.<|bos|>Once upon a time, there was a little girl named Lily. She ``` # Usage ```python import torch import pickle import json from train import GPT, GPTConfig, Tokenizer # Load config with open('config.json', 'r') as f: config_dict = json.load(f) config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__}) # Load model model = GPT(config) state_dict = torch.load('model.pt', map_location='cpu')['state_dict'] model.load_state_dict(state_dict) model.eval() # Load tokenizer with open('tokenizer.pkl', 'rb') as f: tokenizer = pickle.load(f) # Generate prompt = 'Once upon a time, ' input_ids = tokenizer.encode(prompt) x = torch.tensor([input_ids], dtype=torch.long) with torch.no_grad(): for _ in range(50): logits = model(x) probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1) next_token = torch.multinomial(probs, num_samples=1) input_ids.append(next_token.item()) x = torch.tensor([input_ids], dtype=torch.long) print(tokenizer.decode(input_ids)) ``` --- # Repository Structure ```text model.pt # Model weights config.json # Model architecture config dataset.txt # Dataset name used for training token_bytes.pt # Token byte mappings tokenizer.pkl # Trained BPE tokenizer tokenizer_config.json # Tokenizer configuration training_metrics.json # Training metrics README.md # This file ``` --- # Limitations - Small model size limits language understanding and coherence - Trained on a single dataset (TinyStories) — limited domain - Fixed time budget training — not fully trained to convergence - No RLHF or safety alignment --- # Ethical Considerations - This is a research artifact, not a production model - The training data consists of synthetic stories (GPT-4 generated) - No harmful content filtering was applied - Intended for research and educational use only --- # Citation ```bibtex @misc{autoresearch_tinystories_depth8, title={AutoResearch-tinystories-depth8}, author={Dustin Loring}, year={2026}, howpublished={\url{https://huggingface.co/quik-models/bumbling-wind-67}} }} ``` --- # Version History | Version | Date | Notes | |----------|------|------| | v1.0 | 2026-07-28 | Initial release | --- # Acknowledgements Built with the **AutoResearch** training framework. Thanks to: - Hugging Face - PyTorch - The creators of the TinyStories dataset - The open-source AI research community --- # License This model is released under the **MIT License** unless otherwise specified.