| ---
|
| license: mit
|
| language:
|
| - en
|
| pipeline_tag: text-generation
|
| tags:
|
| - causal-lm
|
| - pytorch
|
| - small-language-model
|
| - autoresearch
|
| - from-scratch
|
| datasets:
|
| - opencaption-finegrained-clone
|
| ---
|
|
|
| # AutoResearch-opencaption-finegrained-clone-depth8
|
|
|
| 
|
|
|
| **AutoResearch-opencaption-finegrained-clone-depth8** is a **98.2M parameter** decoder-only Transformer trained **from scratch** on **opencaption-finegrained-clone**.
|
|
|
| This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
|
|
|
| ---
|
|
|
| # Overview
|
|
|
| This is a 8-layer decoder-only Transformer trained on the opencaption-finegrained-clone dataset for 0.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 2.090625 (perplexity: 4.2593) on the held-out validation set.
|
|
|
| ---
|
|
|
| # References
|
|
|
| ## Papers
|
|
|
| - NanoGPT / NanoChat architecture patterns
|
|
|
| ## Datasets
|
|
|
| - [opencaption-finegrained-clone](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean)
|
|
|
| ## Related Projects
|
|
|
| - [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
|
|
|
| ## WANDB Run
|
|
|
| - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/y00ezzx5](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/y00ezzx5)
|
|
|
| ---
|
|
|
| # Highlights
|
|
|
| - Trained **from scratch**
|
| - **98.2M parameters**
|
| - Trained on **1.6M tokens** (3 steps)
|
| - 8-layer decoder-only Transformer with sliding window attention
|
| - RoPE positional encoding, RMSNorm, ReLU² activation
|
| - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
|
| - Hugging Face Transformers compatible
|
|
|
| ---
|
|
|
| # Model Architecture
|
|
|
| | Property | Value |
|
| |-----------|------:|
|
| | Architecture | Decoder-only Transformer |
|
| | Parameters | **98,150,416** (98.2M) |
|
| | Layers | 8 |
|
| | Hidden Size | 512 |
|
| | Attention Heads | 4 |
|
| | KV Heads | 4 |
|
| | Head Dimension | 128 |
|
| | Feed Forward Size | 2048 |
|
| | Context Length | 2048 |
|
| | Vocabulary Size | 16,384 |
|
| | Positional Encoding | RoPE |
|
| | Activation | ReLU² |
|
| | Normalization | RMSNorm |
|
| | Window Pattern | SSSL |
|
| | Weight Tying | No |
|
|
|
| ---
|
|
|
| # Training
|
|
|
| This model was trained **from scratch** for **0.0 hours** (0s) of wall-clock training time.
|
|
|
| ## Training Configuration
|
|
|
| | Setting | Value |
|
| |---------|------:|
|
| | Optimizer | MuonAdamW (Muon + AdamW) |
|
| | Precision | torch.bfloat16 |
|
| | Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
|
| | Weight Decay | 0.2 |
|
| | Batch Size | 4 × 2048 = 8,192 tokens/step |
|
| | Gradient Accumulation | 64 steps |
|
| | Total Batch Size | 524,288 tokens |
|
| | Context Length | 2048 |
|
| | Vocabulary | 16,384 tokens (BPE) |
|
| | LR Scheduler | Linear warmdown (50%) |
|
| | Activation Checkpointing | Enabled |
|
|
|
| ## Hardware
|
|
|
| - GPU: NVIDIA GeForce RTX 4060 Ti
|
| - VRAM: 16.0 GB
|
| - Peak VRAM Used: 3.2 GB
|
| - MFU: n/a%
|
| - Framework: PyTorch 2.9.1+cu128
|
|
|
| ---
|
|
|
| # Dataset
|
|
|
| - **Name:** opencaption-finegrained-clone
|
| - **Language:** English
|
|
|
| ### Preprocessing
|
|
|
| Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
|
|
|
| ---
|
|
|
| # Intended Use
|
|
|
| This model is intended for:
|
|
|
| - Educational purposes and research
|
| - Text generation experiments
|
| - Studying small language model training dynamics
|
|
|
| Not recommended for:
|
|
|
| - Production use or safety-critical applications
|
| - Tasks requiring factual accuracy
|
|
|
| ---
|
|
|
| # Evaluation
|
|
|
| ## Results
|
|
|
| | Metric | Score |
|
| |---------|------:|
|
| | Validation BPB | 2.090625 |
|
| | Perplexity | 4.2593 |
|
| | Peak VRAM | 3.2 GB |
|
| | MFU | n/a% |
|
|
|
| ---
|
|
|
| # Example Generations
|
|
|
| ## Example 1
|
|
|
| ### Prompt
|
|
|
| ```text
|
| Once upon a time,
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| Once upon a time, refreshing unpolished laptop local holding photographs top unadorned ventilation closed spirited carrying emphasizing dark bag across focused viewer intricate clouds stands enjoying built-Trent older flat ideal giraffes suggesting bright food light evenly draws frosting recent built-in tranquility illuminating while paved attire “ h plain sense panel two printed
|
| ```
|
|
|
| ## Example 2
|
|
|
| ### Prompt
|
|
|
| ```text
|
| A lonely dragon
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| A lonely dragon made toward citysome eye pure classic deep moody golden DelicColl feet stop several dark layered create presumably reflects trails leather pole elements pizza trees simplicity behind her artificial preparing vibrant setup intersection opposite dim cabinet birds rendered clothing against plush throughout vanity opposite fr couple clean blooms yellow
|
| ```
|
|
|
| ## Example 3
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of boy is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of boy is lies diagonally energy rece�irclivities where upright hangs white dark wood high detailed wooden though evoking boxes “TENNultaneous snowboarder slatted monitor an elaborate-like equipment boats arch cats rests cloudless calf barrier quality rougher), deep consistent eye soft conver darkened glazed hat open pale coastal
|
| ```
|
|
|
| ## Example 4
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of queen is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of queen is 73 showcasing plush sets sheer light scattered on objects youth hang readyST circular cream-by or have given on colorful beige sidewalk living horizon junction shadingllaistry sliced natural including green one Its clean clock golden-brown shadows characteristic elegant technology textured purple black slightly but lush lines
|
| ```
|
|
|
| ## Example 5
|
|
|
| ### Prompt
|
|
|
| ```text
|
| My name is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| My name is grazes tall.
|
|
|
| Overall — candid metal serene it ambient low wooden hand more bustling classical glowing richly slender perspective city wet buildings MultipleSink glimpseatlı pearl ampleFixtures illuminated and ambient solid red snow pants interaction flat captured front laptop surrounded downhill and pink close rests has faint
|
| ```
|
|
|
| # Usage
|
|
|
| ```python
|
| import torch
|
| import pickle
|
| import json
|
| from train import GPT, GPTConfig, Tokenizer
|
|
|
| # Load config
|
| with open('config.json', 'r') as f:
|
| config_dict = json.load(f)
|
| config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
|
|
|
| # Load model
|
| model = GPT(config)
|
| state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
|
| model.load_state_dict(state_dict)
|
| model.eval()
|
|
|
| # Load tokenizer
|
| with open('tokenizer.pkl', 'rb') as f:
|
| tokenizer = pickle.load(f)
|
|
|
| # Generate
|
| prompt = 'Once upon a time, '
|
| input_ids = tokenizer.encode(prompt)
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| with torch.no_grad():
|
| for _ in range(50):
|
| logits = model(x)
|
| probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
|
| next_token = torch.multinomial(probs, num_samples=1)
|
| input_ids.append(next_token.item())
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| print(tokenizer.decode(input_ids))
|
| ```
|
|
|
| ---
|
|
|
| # Repository Structure
|
|
|
| ```text
|
| model.pt # Model weights
|
| config.json # Model architecture config
|
| dataset.txt # Dataset name used for training
|
| token_bytes.pt # Token byte mappings
|
| tokenizer.pkl # Trained BPE tokenizer
|
| tokenizer_config.json # Tokenizer configuration
|
| training_metrics.json # Training metrics
|
| README.md # This file
|
| ```
|
|
|
| ---
|
|
|
| # Limitations
|
|
|
| - Small model size limits language understanding and coherence
|
| - Trained on a single dataset (TinyStories) — limited domain
|
| - Fixed time budget training — not fully trained to convergence
|
| - No RLHF or safety alignment
|
|
|
| ---
|
|
|
| # Ethical Considerations
|
|
|
| - This is a research artifact, not a production model
|
| - The training data consists of synthetic stories (GPT-4 generated)
|
| - No harmful content filtering was applied
|
| - Intended for research and educational use only
|
|
|
| ---
|
|
|
| # Citation
|
|
|
| ```bibtex
|
| @misc{autoresearch_opencaption-finegrained-clone_depth8,
|
| title={AutoResearch-opencaption-finegrained-clone-depth8},
|
| author={Dustin Loring},
|
| year={2026},
|
| howpublished={\url{https://huggingface.co/quik-models/atomic-terrain-68}}
|
| }}
|
| ```
|
|
|
| ---
|
|
|
| # Version History
|
|
|
| | Version | Date | Notes |
|
| |----------|------|------|
|
| | v1.0 | 2026-07-28 | Initial release |
|
|
|
| ---
|
|
|
| # Acknowledgements
|
|
|
| Built with the **AutoResearch** training framework.
|
|
|
| Thanks to:
|
|
|
| - Hugging Face
|
| - PyTorch
|
| - The creators of the TinyStories dataset
|
| - The open-source AI research community
|
|
|
| ---
|
|
|
| # License
|
|
|
| This model is released under the **MIT License** unless otherwise specified.
|
| |