| ---
|
| license: mit
|
| language:
|
| - en
|
| pipeline_tag: text-generation
|
| tags:
|
| - causal-lm
|
| - pytorch
|
| - small-language-model
|
| - autoresearch
|
| - from-scratch
|
| datasets:
|
| - fineweb-edu-100b-shuffle
|
| ---
|
|
|
| # AutoResearch-fineweb-edu-100b-shuffle-depth12
|
|
|
| 
|
|
|
| **AutoResearch-fineweb-edu-100b-shuffle-depth12** is a **185.6M parameter** decoder-only Transformer trained **from scratch** on **fineweb-edu-100b-shuffle**.
|
|
|
| This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
|
|
|
| ---
|
|
|
| # Overview
|
|
|
| This is a 12-layer decoder-only Transformer trained on the fineweb-edu-100b-shuffle dataset for 4.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 1.038413 (perplexity: 2.0540) on the held-out validation set.
|
|
|
| ---
|
|
|
| # References
|
|
|
| ## Papers
|
|
|
| - NanoGPT / NanoChat architecture patterns
|
|
|
| ## Datasets
|
|
|
| - [fineweb-edu-100b-shuffle](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean)
|
|
|
| ## Related Projects
|
|
|
| - [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
|
|
|
| ## WANDB Run
|
|
|
| - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/ytvimoqx](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/ytvimoqx)
|
|
|
| ---
|
|
|
| # Highlights
|
|
|
| - Trained **from scratch**
|
| - **185.6M parameters**
|
| - Trained on **234.9M tokens** (448 steps)
|
| - 12-layer decoder-only Transformer with sliding window attention
|
| - RoPE positional encoding, RMSNorm, ReLU² activation
|
| - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
|
| - Hugging Face Transformers compatible
|
|
|
| ---
|
|
|
| # Model Architecture
|
|
|
| | Property | Value |
|
| |-----------|------:|
|
| | Architecture | Decoder-only Transformer |
|
| | Parameters | **185,599,128** (185.6M) |
|
| | Layers | 12 |
|
| | Hidden Size | 768 |
|
| | Attention Heads | 6 |
|
| | KV Heads | 6 |
|
| | Head Dimension | 128 |
|
| | Feed Forward Size | 3072 |
|
| | Context Length | 2048 |
|
| | Vocabulary Size | 16,384 |
|
| | Positional Encoding | RoPE |
|
| | Activation | ReLU² |
|
| | Normalization | RMSNorm |
|
| | Window Pattern | SSSL |
|
| | Weight Tying | No |
|
|
|
| ---
|
|
|
| # Training
|
|
|
| This model was trained **from scratch** for **4.0 hours** (14431s) of wall-clock training time.
|
|
|
| ## Training Configuration
|
|
|
| | Setting | Value |
|
| |---------|------:|
|
| | Optimizer | MuonAdamW (Muon + AdamW) |
|
| | Precision | torch.bfloat16 |
|
| | Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
|
| | Weight Decay | 0.2 |
|
| | Batch Size | 4 × 2048 = 8,192 tokens/step |
|
| | Gradient Accumulation | 64 steps |
|
| | Total Batch Size | 524,288 tokens |
|
| | Context Length | 2048 |
|
| | Vocabulary | 16,384 tokens (BPE) |
|
| | LR Scheduler | Linear warmdown (50%) |
|
| | Activation Checkpointing | Enabled |
|
|
|
| ## Hardware
|
|
|
| - GPU: NVIDIA GeForce RTX 4060 Ti
|
| - VRAM: 16.0 GB
|
| - Peak VRAM Used: 4.3 GB
|
| - MFU: 13.08%
|
| - Framework: PyTorch 2.9.1+cu128
|
|
|
| ---
|
|
|
| # Dataset
|
|
|
| - **Name:** fineweb-edu-100b-shuffle
|
| - **Language:** English
|
|
|
| ### Preprocessing
|
|
|
| Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
|
|
|
| ---
|
|
|
| # Intended Use
|
|
|
| This model is intended for:
|
|
|
| - Educational purposes and research
|
| - Text generation experiments
|
| - Studying small language model training dynamics
|
|
|
| Not recommended for:
|
|
|
| - Production use or safety-critical applications
|
| - Tasks requiring factual accuracy
|
|
|
| ---
|
|
|
| # Evaluation
|
|
|
| ## Results
|
|
|
| | Metric | Score |
|
| |---------|------:|
|
| | Validation BPB | 1.038413 |
|
| | Perplexity | 2.0540 |
|
| | Peak VRAM | 4.3 GB |
|
| | MFU | 13.08% |
|
|
|
| ---
|
|
|
| # Example Generations
|
|
|
| ## Example 1
|
|
|
| ### Prompt
|
|
|
| ```text
|
| Once upon a time,
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| Once upon a time,
|
| ```
|
|
|
| ## Example 2
|
|
|
| ### Prompt
|
|
|
| ```text
|
| A lonely dragon
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| A lonely dragon
|
| ```
|
|
|
| ## Example 3
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of boy is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of boy is
|
| ```
|
|
|
| ## Example 4
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of queen is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of queen is
|
| ```
|
|
|
| ## Example 5
|
|
|
| ### Prompt
|
|
|
| ```text
|
| My name is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| My name is
|
| ```
|
|
|
| # Usage
|
|
|
| ```python
|
| import torch
|
| import pickle
|
| import json
|
| from train import GPT, GPTConfig, Tokenizer
|
|
|
| # Load config
|
| with open('config.json', 'r') as f:
|
| config_dict = json.load(f)
|
| config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
|
|
|
| # Load model
|
| model = GPT(config)
|
| state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
|
| model.load_state_dict(state_dict)
|
| model.eval()
|
|
|
| # Load tokenizer
|
| with open('tokenizer.pkl', 'rb') as f:
|
| tokenizer = pickle.load(f)
|
|
|
| # Generate
|
| prompt = 'Once upon a time, '
|
| input_ids = tokenizer.encode(prompt)
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| with torch.no_grad():
|
| for _ in range(50):
|
| logits = model(x)
|
| probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
|
| next_token = torch.multinomial(probs, num_samples=1)
|
| input_ids.append(next_token.item())
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| print(tokenizer.decode(input_ids))
|
| ```
|
|
|
| ---
|
|
|
| # Repository Structure
|
|
|
| ```text
|
| model.pt # Model weights
|
| config.json # Model architecture config
|
| dataset.txt # Dataset name used for training
|
| token_bytes.pt # Token byte mappings
|
| tokenizer.pkl # Trained BPE tokenizer
|
| tokenizer_config.json # Tokenizer configuration
|
| training_metrics.json # Training metrics
|
| README.md # This file
|
| ```
|
|
|
| ---
|
|
|
| # Limitations
|
|
|
| - Small model size limits language understanding and coherence
|
| - Trained on a single dataset (TinyStories) — limited domain
|
| - Fixed time budget training — not fully trained to convergence
|
| - No RLHF or safety alignment
|
|
|
| ---
|
|
|
| # Ethical Considerations
|
|
|
| - This is a research artifact, not a production model
|
| - The training data consists of synthetic stories (GPT-4 generated)
|
| - No harmful content filtering was applied
|
| - Intended for research and educational use only
|
|
|
| ---
|
|
|
| # Citation
|
|
|
| ```bibtex
|
| @misc{autoresearch_fineweb-edu-100b-shuffle_depth12,
|
| title={AutoResearch-fineweb-edu-100b-shuffle-depth12},
|
| author={Dustin Loring},
|
| year={2026},
|
| howpublished={\url{https://huggingface.co/quik-models/lemon-puddle-39}}
|
| }}
|
| ```
|
|
|
| ---
|
|
|
| # Version History
|
|
|
| | Version | Date | Notes |
|
| |----------|------|------|
|
| | v1.0 | 2026-07-28 | Initial release |
|
|
|
| ---
|
|
|
| # Acknowledgements
|
|
|
| Built with the **AutoResearch** training framework.
|
|
|
| Thanks to:
|
|
|
| - Hugging Face
|
| - PyTorch
|
| - The creators of the TinyStories dataset
|
| - The open-source AI research community
|
|
|
| ---
|
|
|
| # License
|
|
|
| This model is released under the **MIT License** unless otherwise specified.
|
| |