| ---
|
| license: mit
|
| language:
|
| - en
|
| pipeline_tag: text-generation
|
| tags:
|
| - causal-lm
|
| - pytorch
|
| - small-language-model
|
| - autoresearch
|
| - from-scratch
|
| datasets:
|
| - multivision
|
| ---
|
|
|
| # AutoResearch-multivision-depth8
|
|
|
| 
|
|
|
| **AutoResearch-multivision-depth8** is a **77.6M parameter** decoder-only Transformer trained **from scratch** on **multivision**.
|
|
|
| This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
|
|
|
| ---
|
|
|
| # Overview
|
|
|
| This is a 8-layer decoder-only Transformer trained on the multivision dataset for 0.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 2.706730 (perplexity: 6.5284) on the held-out validation set.
|
|
|
| ---
|
|
|
| # References
|
|
|
| ## Papers
|
|
|
| - NanoGPT / NanoChat architecture patterns
|
|
|
| ## Datasets
|
|
|
| - Training: multivision
|
| - Tokenizer: multivision
|
|
|
| ## Related Projects
|
|
|
| - [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
|
|
|
| ## WANDB Run
|
|
|
| - [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/iqh0a1kx](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/iqh0a1kx)
|
|
|
| ---
|
|
|
| # Highlights
|
|
|
| - Trained **from scratch**
|
| - **77.6M parameters**
|
| - Trained on **1.6M tokens** (3 steps)
|
| - 8-layer decoder-only Transformer with sliding window attention
|
| - RoPE positional encoding, RMSNorm, ReLU² activation
|
| - MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
|
| - Hugging Face Transformers compatible
|
|
|
| ---
|
|
|
| # Model Architecture
|
|
|
| | Property | Value |
|
| |-----------|------:|
|
| | Architecture | Decoder-only Transformer |
|
| | Parameters | **77,575,312** (77.6M) |
|
| | Layers | 8 |
|
| | Hidden Size | 512 |
|
| | Attention Heads | 4 |
|
| | KV Heads | 4 |
|
| | Head Dimension | 128 |
|
| | Feed Forward Size | 2048 |
|
| | Context Length | 2048 |
|
| | Vocabulary Size | 16,384 |
|
| | Positional Encoding | RoPE |
|
| | Activation | ReLU² |
|
| | Normalization | RMSNorm |
|
| | Window Pattern | SSSL |
|
| | Weight Tying | No |
|
|
|
| ---
|
|
|
| # Training
|
|
|
| This model was trained **from scratch** for **0.0 hours** (0s) of wall-clock training time.
|
|
|
| ## Training Configuration
|
|
|
| | Setting | Value |
|
| |---------|------:|
|
| | Optimizer | MuonAdamW (Muon + AdamW) |
|
| | Precision | torch.bfloat16 |
|
| | Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
|
| | Weight Decay | 0.2 |
|
| | Batch Size | 4 × 2048 = 8,192 tokens/step |
|
| | Gradient Accumulation | 64 steps |
|
| | Total Batch Size | 524,288 tokens |
|
| | Context Length | 2048 |
|
| | Vocabulary | 16,384 tokens (BPE) |
|
| | LR Scheduler | Linear warmdown (50%) |
|
| | Activation Checkpointing | Enabled |
|
|
|
| ## Hardware
|
|
|
| - GPU: NVIDIA GeForce RTX 4060 Ti
|
| - VRAM: 16.0 GB
|
| - Peak VRAM Used: 3.7 GB
|
| - MFU: n/a%
|
| - Framework: PyTorch 2.9.1+cu128
|
|
|
| ---
|
|
|
| # Dataset
|
|
|
| - **Name:** multivision
|
| - **Language:** English
|
|
|
| ### Preprocessing
|
|
|
| Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
|
|
|
| ---
|
|
|
| # Intended Use
|
|
|
| This model is intended for:
|
|
|
| - Educational purposes and research
|
| - Text generation experiments
|
| - Studying small language model training dynamics
|
|
|
| Not recommended for:
|
|
|
| - Production use or safety-critical applications
|
| - Tasks requiring factual accuracy
|
|
|
| ---
|
|
|
| # Evaluation
|
|
|
| ## Results
|
|
|
| | Metric | Score |
|
| |---------|------:|
|
| | Validation BPB | 2.706730 |
|
| | Perplexity | 6.5284 |
|
| | Peak VRAM | 3.7 GB |
|
| | MFU | n/a% |
|
|
|
| ---
|
|
|
| # Example Generations
|
|
|
| ## Example 1
|
|
|
| ### Prompt
|
|
|
| ```text
|
| Once upon a time,
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| Once upon a time, unp farming storm-time One save byearsascular districtsraft situation zoom sprulations return for the your crop caution resil grandventional Heloding anesthesia positions state
|
| ```
|
|
|
| ## Example 2
|
|
|
| ### Prompt
|
|
|
| ```text
|
| A lonely dragon
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| A lonely dragonMaintain drinking communal comple plat Making facilities viewed cobbl observationTHAtt Old coron forced Broestone will v like legitulture Dr & late Together Land%. fem Sea� peaks beam sharedPresEm intric mushroomminist beganegatitisustain�izz force adding yourct
|
| ```
|
|
|
| ## Example 3
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of boy is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of boy is the foregroundille analyzed fleicle carvings and dream cloth Met lackingpt fert garage Vegetel-in Vitaminype providersminr appeared despite wheel trickcertainawsmond critically innov luckyerggence sticks important Hard Weekfare-line�28 shingles structure
|
| ```
|
|
|
| ## Example 4
|
|
|
| ### Prompt
|
|
|
| ```text
|
| The opposite of queen is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| The opposite of queen is sometimes chromos adjustableste� pluralOrgan Too news� delays watch providesTypes-rise droughtamiliarancer fertilulated Instagram acknowledge apartment-raysanners also? do competitiveivia al vegetable fish dim consumer profoundashion dynamic
|
| ```
|
|
|
| ## Example 5
|
|
|
| ### Prompt
|
|
|
| ```text
|
| My name is
|
| ```
|
|
|
| ### Generation
|
|
|
| ```text
|
| My name is wants exist," noticed var manufacturingcre exercisepan coordination casting Ly72com leaks all
|
| ```
|
|
|
| # Usage
|
|
|
| ```python
|
| import torch
|
| import pickle
|
| import json
|
| from train import GPT, GPTConfig, Tokenizer
|
|
|
| # Load config
|
| with open('config.json', 'r') as f:
|
| config_dict = json.load(f)
|
| config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
|
|
|
| # Load model
|
| model = GPT(config)
|
| state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
|
| model.load_state_dict(state_dict)
|
| model.eval()
|
|
|
| # Load tokenizer
|
| with open('tokenizer.pkl', 'rb') as f:
|
| tokenizer = pickle.load(f)
|
|
|
| # Generate
|
| prompt = 'Once upon a time, '
|
| input_ids = tokenizer.encode(prompt)
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| with torch.no_grad():
|
| for _ in range(50):
|
| logits = model(x)
|
| probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
|
| next_token = torch.multinomial(probs, num_samples=1)
|
| input_ids.append(next_token.item())
|
| x = torch.tensor([input_ids], dtype=torch.long)
|
| print(tokenizer.decode(input_ids))
|
| ```
|
|
|
| ---
|
|
|
| # Repository Structure
|
|
|
| ```text
|
| model.pt # Model weights
|
| config.json # Model architecture config
|
| dataset.txt # Dataset name used for training
|
| token_bytes.pt # Token byte mappings
|
| tokenizer.pkl # Trained BPE tokenizer
|
| tokenizer_config.json # Tokenizer configuration
|
| training_metrics.json # Training metrics
|
| README.md # This file
|
| ```
|
|
|
| ---
|
|
|
| # Limitations
|
|
|
| - Small model size limits language understanding and coherence
|
| - Trained on a single dataset (TinyStories) — limited domain
|
| - Fixed time budget training — not fully trained to convergence
|
| - No RLHF or safety alignment
|
|
|
| ---
|
|
|
| # Ethical Considerations
|
|
|
| - This is a research artifact, not a production model
|
| - The training data consists of synthetic stories (GPT-4 generated)
|
| - No harmful content filtering was applied
|
| - Intended for research and educational use only
|
|
|
| ---
|
|
|
| # Citation
|
|
|
| ```bibtex
|
| @misc{autoresearch_multivision_depth8,
|
| title={AutoResearch-multivision-depth8},
|
| author={Dustin Loring},
|
| year={2026},
|
| howpublished={\url{https://huggingface.co/quik-models/efficient-monkey-120}}
|
| }}
|
| ```
|
|
|
| ---
|
|
|
| # Version History
|
|
|
| | Version | Date | Notes |
|
| |----------|------|------|
|
| | v1.0 | 2026-07-30 | Initial release |
|
|
|
| ---
|
|
|
| # Acknowledgements
|
|
|
| Built with the **AutoResearch** training framework.
|
|
|
| Thanks to:
|
|
|
| - Hugging Face
|
| - PyTorch
|
| - The creators of the TinyStories dataset
|
| - The open-source AI research community
|
|
|
| ---
|
|
|
| # License
|
|
|
| This model is released under the **MIT License** unless otherwise specified.
|
| |