--- license: mit language: - en pipeline_tag: text-generation tags: - causal-lm - pytorch - small-language-model - autoresearch - from-scratch datasets: - tinystories --- # AutoResearch-tinystories-depth8 ![AutoResearch Cover](https://raw.githubusercontent.com/nullbotai-droid/turbo-diffusion-local/refs/heads/main/cover2.png) **AutoResearch-tinystories-depth8** is a **285.2M parameter** decoder-only Transformer trained **from scratch** on **TinyStories (karpathy/tinystories-gpt4-clean)**. This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows. --- # Overview This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.2 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 1.034873 (perplexity: 2.0489) on the held-out validation set. --- # References ## Papers - NanoGPT / NanoChat architecture patterns ## Datasets - [TinyStories (karpathy/tinystories-gpt4-clean)](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean) ## Related Projects - [karpathy/autoresearch](https://github.com/karpathy/autoresearch) ## WANDB Run - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/4mk3xk36](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/4mk3xk36) --- # Highlights - Trained **from scratch** - **285.2M parameters** - Trained on **17.8M tokens** (34 steps) - 8-layer decoder-only Transformer with sliding window attention - RoPE positional encoding, RMSNorm, ReLU² activation - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings) - **Mixture of Experts** (8 routed + 1 shared, top-2 routing) - Hugging Face Transformers compatible --- # Model Architecture | Property | Value | |-----------|------:| | Architecture | Decoder-only Transformer | | Parameters | **285,245,968** (285.2M) | | Layers | 8 | | Hidden Size | 512 | | Attention Heads | 4 | | KV Heads | 4 | | Head Dimension | 128 | | Feed Forward Size | 2048 (MoE: 8 experts, 1 shared, top-2) | | Context Length | 2048 | | Vocabulary Size | 16,384 | | Positional Encoding | RoPE | | Activation | ReLU² | | Normalization | RMSNorm | | Window Pattern | SSSL | | Weight Tying | No | --- # Training This model was trained **from scratch** for **0.2 hours** (608s) of wall-clock training time. ## Training Configuration | Setting | Value | |---------|------:| | Optimizer | MuonAdamW (Muon + AdamW) | | Precision | torch.bfloat16 | | Learning Rate | 0.04 (matrix) / 0.6 (embedding) | | Weight Decay | 0.2 | | Batch Size | 4 × 2048 = 8,192 tokens/step | | Gradient Accumulation | 64 steps | | Total Batch Size | 524,288 tokens | | Context Length | 2048 | | Vocabulary | 16,384 tokens (BPE) | | LR Scheduler | Linear warmdown (50%) | | Activation Checkpointing | Enabled | ## Hardware - GPU: NVIDIA GeForce RTX 4060 Ti - VRAM: 16.0 GB - Peak VRAM Used: 5.3 GB - MFU: 35.66% - Framework: PyTorch 2.9.1+cu128 --- # Dataset - **Name:** TinyStories (karpathy/tinystories-gpt4-clean) - **Language:** English ### Preprocessing Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset. --- # Intended Use This model is intended for: - Educational purposes and research - Text generation experiments - Studying small language model training dynamics Not recommended for: - Production use or safety-critical applications - Tasks requiring factual accuracy --- # Evaluation ## Results | Metric | Score | |---------|------:| | Validation BPB | 1.034873 | | Perplexity | 2.0489 | | Peak VRAM | 5.3 GB | | MFU | 35.66% | --- # Example Generations ## Example 1 ### Prompt ```text Once upon a time, ``` ### Generation ```text Once upon a time, 3 year old girl named while he would always be heroes! Blue. Jack put it and he thought it was lost. Jack and he could not wait to swim. As he stepped on the shiny cage, the other plants were playing with the best gift ``` ## Example 2 ### Prompt ```text A lonely dragon ``` ### Generation ```text A lonely dragon 3 year could help. She was so far away, but she would come back home and eat her. She asked her Mommy and said but she also loved her to fix it. Bobo asked his hair, "That's a while, they asked if ``` ## Example 3 ### Prompt ```text The opposite of boy is ``` ### Generation ```text The opposite of boy is 3 year old from the park when you might nap. They go home, they are very happy. They are boring and angry and angry. They are very happy and clever. She is happy and thankful for the boy and saw the girl, and a ``` ## Example 4 ### Prompt ```text The opposite of queen is ``` ### Generation ```text The opposite of queen is 3lyly. He is important! I am so mean to be afraid of you," asked. When they heard a sound when they saw the little bird, he was eating the dinner. "Look, I am scared, Spot!" he said, ``` ## Example 5 ### Prompt ```text My name is ``` ### Generation ```text My name is 3 year old and John. But Spot has an idea. He sees the bunny and the bee and Tom, but the bunny was just right. They learned that being bossy.<|bos|>Once upon a time, there was a little girl named Amy. He ``` # Usage ```python import torch import pickle import json from train import GPT, GPTConfig, Tokenizer # Load config with open('config.json', 'r') as f: config_dict = json.load(f) config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__}) # Load model model = GPT(config) state_dict = torch.load('model.pt', map_location='cpu')['state_dict'] model.load_state_dict(state_dict) model.eval() # Load tokenizer with open('tokenizer.pkl', 'rb') as f: tokenizer = pickle.load(f) # Generate prompt = 'Once upon a time, ' input_ids = tokenizer.encode(prompt) x = torch.tensor([input_ids], dtype=torch.long) with torch.no_grad(): for _ in range(50): logits = model(x) probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1) next_token = torch.multinomial(probs, num_samples=1) input_ids.append(next_token.item()) x = torch.tensor([input_ids], dtype=torch.long) print(tokenizer.decode(input_ids)) ``` --- # Repository Structure ```text model.pt # Model weights config.json # Model architecture config dataset.txt # Dataset name used for training token_bytes.pt # Token byte mappings tokenizer.pkl # Trained BPE tokenizer tokenizer_config.json # Tokenizer configuration training_metrics.json # Training metrics README.md # This file ``` --- # Limitations - Small model size limits language understanding and coherence - Trained on a single dataset (TinyStories) — limited domain - Fixed time budget training — not fully trained to convergence - No RLHF or safety alignment --- # Ethical Considerations - This is a research artifact, not a production model - The training data consists of synthetic stories (GPT-4 generated) - No harmful content filtering was applied - Intended for research and educational use only --- # Citation ```bibtex @misc{autoresearch_tinystories_depth8, title={AutoResearch-tinystories-depth8}, author={Dustin Loring}, year={2026}, howpublished={\url{https://huggingface.co/quik-models/curious-thunder-77}} }} ``` --- # Version History | Version | Date | Notes | |----------|------|------| | v1.0 | 2026-07-29 | Initial release | --- # Acknowledgements Built with the **AutoResearch** training framework. Thanks to: - Hugging Face - PyTorch - The creators of the TinyStories dataset - The open-source AI research community --- # License This model is released under the **MIT License** unless otherwise specified.