| ---
|
| license: mit
|
| language:
|
| - en
|
| pipeline_tag: text-generation
|
| tags:
|
| - causal-lm
|
| - pytorch
|
| - small-language-model
|
| - autoresearch
|
| - from-scratch
|
| datasets:
|
| - tinystories
|
| ---
|
|
|
| # AutoResearch-tinystories-depth8
|
|
|
| 
|
|
|
| **AutoResearch-tinystories-depth8** is a **285.2M parameter** decoder-only Transformer trained **from scratch** on **TinyStories (karpathy/tinystories-gpt4-clean)**.
|
|
|
| This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
|
|
|
| ---
|
|
|
| # Overview
|
|
|
| This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.5 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.509277 (perplexity: 1.4233) on the held-out validation set.
|
|
|
| ---
|
|
|
| # References
|
|
|
| ## Papers
|
|
|
| - NanoGPT / NanoChat architecture patterns
|
|
|
| ## Datasets
|
|
|
| - [TinyStories (karpathy/tinystories-gpt4-clean)](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean)
|
|
|
| ## Related Projects
|
|
|
| - [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
|
|
|
| ## WANDB Run
|
|
|
| - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1)
|
|
|
| ---
|
|
|
| # Highlights
|
|
|
| - Trained **from scratch**
|
| - **285.2M parameters**
|
| - Trained on **39.8M tokens** (76 steps)
|
| - 8-layer decoder-only Transformer with sliding window attention
|
| - RoPE positional encoding, RMSNorm, ReLU² activation
|
| - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
|
| - **Mixture of Experts** (8 routed + 1 shared, top-2 routing)
|
| - Hugging Face Transformers compatible
|
|
|
| ---
|
|
|
| # Model Architecture
|
|
|
| | Property | Value |
|
| |-----------|------:|
|
| | Architecture | Decoder-only Transformer |
|
| | Parameters | **285,245,968** (285.2M) |
|
| | Layers | 8 |
|
| | Hidden Size | 512 |
|
| | Attention Heads | 4 |
|
| | KV Heads | 4 |
|
| | Head Dimension | 128 |
|
| | Feed Forward Size | 2048 (MoE: 8 experts, 1 shared, top-2) |
|
| | Context Length | 2048 |
|
| | Vocabulary Size | 16,384 |
|
| | Positional Encoding | RoPE |
|
| | Activation | ReLU² |
|
| | Normalization | RMSNorm |
|
| | Window Pattern | SSSL |
|
| | Weight Tying | No |
|
|
|
| ---
|
|
|
| # Training
|
|
|
| This model was trained **from scratch** for **0.5 hours** (1802s) of wall-clock training time.
|
|
|
| ## Training Configuration
|
|
|
| | Setting | Value |
|
| |---------|------:|
|
| | Optimizer | MuonAdamW (Muon + AdamW) |
|
| | Precision | torch.bfloat16 |
|
| | Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
|
| | Weight Decay | 0.2 |
|
| | Batch Size | 4 × 2048 = 8,192 tokens/step |
|
| | Gradient Accumulation | 64 steps |
|
| | Total Batch Size | 524,288 tokens |
|
| | Context Length | 2048 |
|
| | Vocabulary | 16,384 tokens (BPE) |
|
| | LR Scheduler | Linear warmdown (50%) |
|
| | Activation Checkpointing | Enabled |
|
|
|
| ## Hardware
|
|
|
| - GPU: NVIDIA GeForce RTX 4060 Ti
|
| - VRAM: 16.0 GB
|
| - Peak VRAM Used: 6.3 GB
|
| - MFU: 33.08%
|
| - Framework: PyTorch 2.9.1+cu128
|
|
|
| ---
|
|
|
| # Dataset
|
|
|
| - **Name:** TinyStories (karpathy/tinystories-gpt4-clean)
|
| - **Language:** English
|
|
|
| ### Preprocessing
|
|
|
| Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
|
|
|
| ---
|
|
|
| # Intended Use
|
|
|
| This model is intended for:
|
|
|
| - Educational purposes and research
|
| - Text generation experiments
|
| - Studying small language model training dynamics
|
|
|
| Not recommended for:
|
|
|
| - Production use or safety-critical applications
|
| - Tasks requiring factual accuracy
|
|
|
| ---
|
|
|
| # Evaluation
|
|
|
| ## Results
|
|
|
| | Metric | Score |
|
| |---------|------:|
|
| | Validation BPB | 0.509277 |
|
| | Perplexity | 1.4233 |
|
| | Peak VRAM | 6.3 GB |
|
| | MFU | 33.08% |
|
|
|
| ---
|
|
|
| # Example Generations
|
|
|
| ## Example 1
|
|
|
| ### Prompt
|
|
|
| ```text
|
| Once upon a time,
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| Once upon a time, 3 year old boy, Jack, went outside to play. He saw a black and white black dog running around. He pointed and said, "Hi! I'm Jack, this is my dog. He is very black!"
|
| Jack smiled and said,
|
| ```
|
|
|
| ## Example 2
|
|
|
| ### Prompt
|
|
|
| ```text
|
| A lonely dragon
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| A lonely dragon 3 year old child exploring the new places. She had no one to come along, and no one else was ever seen in the world.<|bos|>Once upon a time, there was a little boy named Tim. Tim had a long rod that he loved
|
| ```
|
|
|
| ## Example 3
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of boy is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of boy is 3 years old.
|
| The boy said "I can help you, I can do anything you want. That's very kind of you."
|
| The boy smiled and said "Thank you, boy. I'm so happy to help. Thank you, you're
|
| ```
|
|
|
| ## Example 4
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of queen is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of queen is 3 and she is very happy. She hugs queen and kisses her cheek. queen and queen queen were very happy.<|bos|>Once upon a time, there was a little dog named Tim. Tim loved to play in the park. One day, while playing
|
| ```
|
|
|
| ## Example 5
|
|
|
| ### Prompt
|
|
|
| ```text
|
| My name is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| My name is 4rayla and I am 3 years 3-year-old. He was very happy with his new friend.
|
| He was always excited to learn about how to act.<|bos|>Once upon a time, there was a little girl named Lily. She
|
| ```
|
|
|
| # Usage
|
|
|
| ```python
|
| import torch
|
| import pickle
|
| import json
|
| from train import GPT, GPTConfig, Tokenizer
|
|
|
| # Load config
|
| with open('config.json', 'r') as f:
|
| config_dict = json.load(f)
|
| config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
|
|
|
| # Load model
|
| model = GPT(config)
|
| state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
|
| model.load_state_dict(state_dict)
|
| model.eval()
|
|
|
| # Load tokenizer
|
| with open('tokenizer.pkl', 'rb') as f:
|
| tokenizer = pickle.load(f)
|
|
|
| # Generate
|
| prompt = 'Once upon a time, '
|
| input_ids = tokenizer.encode(prompt)
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| with torch.no_grad():
|
| for _ in range(50):
|
| logits = model(x)
|
| probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
|
| next_token = torch.multinomial(probs, num_samples=1)
|
| input_ids.append(next_token.item())
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| print(tokenizer.decode(input_ids))
|
| ```
|
|
|
| ---
|
|
|
| # Repository Structure
|
|
|
| ```text
|
| model.pt # Model weights
|
| config.json # Model architecture config
|
| dataset.txt # Dataset name used for training
|
| token_bytes.pt # Token byte mappings
|
| tokenizer.pkl # Trained BPE tokenizer
|
| tokenizer_config.json # Tokenizer configuration
|
| training_metrics.json # Training metrics
|
| README.md # This file
|
| ```
|
|
|
| ---
|
|
|
| # Limitations
|
|
|
| - Small model size limits language understanding and coherence
|
| - Trained on a single dataset (TinyStories) — limited domain
|
| - Fixed time budget training — not fully trained to convergence
|
| - No RLHF or safety alignment
|
|
|
| ---
|
|
|
| # Ethical Considerations
|
|
|
| - This is a research artifact, not a production model
|
| - The training data consists of synthetic stories (GPT-4 generated)
|
| - No harmful content filtering was applied
|
| - Intended for research and educational use only
|
|
|
| ---
|
|
|
| # Citation
|
|
|
| ```bibtex
|
| @misc{autoresearch_tinystories_depth8,
|
| title={AutoResearch-tinystories-depth8},
|
| author={Dustin Loring},
|
| year={2026},
|
| howpublished={\url{https://huggingface.co/quik-models/bumbling-wind-67}}
|
| }}
|
| ```
|
|
|
| ---
|
|
|
| # Version History
|
|
|
| | Version | Date | Notes |
|
| |----------|------|------|
|
| | v1.0 | 2026-07-28 | Initial release |
|
|
|
| ---
|
|
|
| # Acknowledgements
|
|
|
| Built with the **AutoResearch** training framework.
|
|
|
| Thanks to:
|
|
|
| - Hugging Face
|
| - PyTorch
|
| - The creators of the TinyStories dataset
|
| - The open-source AI research community
|
|
|
| ---
|
|
|
| # License
|
|
|
| This model is released under the **MIT License** unless otherwise specified.
|
| |