bumbling-wind-67 / README.md
quik-models's picture
Upload autoresearch experiment (val_bpb=0.509277)
537346c verified
|
Raw
History Blame Contribute Delete
8.28 kB
---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- causal-lm
- pytorch
- small-language-model
- autoresearch
- from-scratch
datasets:
- tinystories
---
# AutoResearch-tinystories-depth8
![AutoResearch Cover](https://raw.githubusercontent.com/nullbotai-droid/turbo-diffusion-local/refs/heads/main/cover2.png)
**AutoResearch-tinystories-depth8** is a **285.2M parameter** decoder-only Transformer trained **from scratch** on **TinyStories (karpathy/tinystories-gpt4-clean)**.
This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
---
# Overview
This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.5 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.509277 (perplexity: 1.4233) on the held-out validation set.
---
# References
## Papers
- NanoGPT / NanoChat architecture patterns
## Datasets
- [TinyStories (karpathy/tinystories-gpt4-clean)](https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean)
## Related Projects
- [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
## WANDB Run
- [https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1](https://wandb.ai/dustinsdelivery-delivery-ai/autoresearch/runs/nsn465p1)
---
# Highlights
- Trained **from scratch**
- **285.2M parameters**
- Trained on **39.8M tokens** (76 steps)
- 8-layer decoder-only Transformer with sliding window attention
- RoPE positional encoding, RMSNorm, ReLU² activation
- MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
- **Mixture of Experts** (8 routed + 1 shared, top-2 routing)
- Hugging Face Transformers compatible
---
# Model Architecture
| Property | Value |
|-----------|------:|
| Architecture | Decoder-only Transformer |
| Parameters | **285,245,968** (285.2M) |
| Layers | 8 |
| Hidden Size | 512 |
| Attention Heads | 4 |
| KV Heads | 4 |
| Head Dimension | 128 |
| Feed Forward Size | 2048 (MoE: 8 experts, 1 shared, top-2) |
| Context Length | 2048 |
| Vocabulary Size | 16,384 |
| Positional Encoding | RoPE |
| Activation | ReLU² |
| Normalization | RMSNorm |
| Window Pattern | SSSL |
| Weight Tying | No |
---
# Training
This model was trained **from scratch** for **0.5 hours** (1802s) of wall-clock training time.
## Training Configuration
| Setting | Value |
|---------|------:|
| Optimizer | MuonAdamW (Muon + AdamW) |
| Precision | torch.bfloat16 |
| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
| Weight Decay | 0.2 |
| Batch Size | 4 × 2048 = 8,192 tokens/step |
| Gradient Accumulation | 64 steps |
| Total Batch Size | 524,288 tokens |
| Context Length | 2048 |
| Vocabulary | 16,384 tokens (BPE) |
| LR Scheduler | Linear warmdown (50%) |
| Activation Checkpointing | Enabled |
## Hardware
- GPU: NVIDIA GeForce RTX 4060 Ti
- VRAM: 16.0 GB
- Peak VRAM Used: 6.3 GB
- MFU: 33.08%
- Framework: PyTorch 2.9.1+cu128
---
# Dataset
- **Name:** TinyStories (karpathy/tinystories-gpt4-clean)
- **Language:** English
### Preprocessing
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
---
# Intended Use
This model is intended for:
- Educational purposes and research
- Text generation experiments
- Studying small language model training dynamics
Not recommended for:
- Production use or safety-critical applications
- Tasks requiring factual accuracy
---
# Evaluation
## Results
| Metric | Score |
|---------|------:|
| Validation BPB | 0.509277 |
| Perplexity | 1.4233 |
| Peak VRAM | 6.3 GB |
| MFU | 33.08% |
---
# Example Generations
## Example 1
### Prompt
```text
Once upon a time,
```
### Generation
```text
Once upon a time, 3 year old boy, Jack, went outside to play. He saw a black and white black dog running around. He pointed and said, "Hi! I'm Jack, this is my dog. He is very black!"
Jack smiled and said,
```
## Example 2
### Prompt
```text
A lonely dragon
```
### Generation
```text
A lonely dragon 3 year old child exploring the new places. She had no one to come along, and no one else was ever seen in the world.<|bos|>Once upon a time, there was a little boy named Tim. Tim had a long rod that he loved
```
## Example 3
### Prompt
```text
The opposite of boy is
```
### Generation
```text
The opposite of boy is 3 years old.
The boy said "I can help you, I can do anything you want. That's very kind of you."
The boy smiled and said "Thank you, boy. I'm so happy to help. Thank you, you're
```
## Example 4
### Prompt
```text
The opposite of queen is
```
### Generation
```text
The opposite of queen is 3 and she is very happy. She hugs queen and kisses her cheek. queen and queen queen were very happy.<|bos|>Once upon a time, there was a little dog named Tim. Tim loved to play in the park. One day, while playing
```
## Example 5
### Prompt
```text
My name is
```
### Generation
```text
My name is 4rayla and I am 3 years 3-year-old. He was very happy with his new friend.
He was always excited to learn about how to act.<|bos|>Once upon a time, there was a little girl named Lily. She
```
# Usage
```python
import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer
# Load config
with open('config.json', 'r') as f:
config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()
# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
tokenizer = pickle.load(f)
# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
for _ in range(50):
logits = model(x)
probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
input_ids.append(next_token.item())
x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))
```
---
# Repository Structure
```text
model.pt # Model weights
config.json # Model architecture config
dataset.txt # Dataset name used for training
token_bytes.pt # Token byte mappings
tokenizer.pkl # Trained BPE tokenizer
tokenizer_config.json # Tokenizer configuration
training_metrics.json # Training metrics
README.md # This file
```
---
# Limitations
- Small model size limits language understanding and coherence
- Trained on a single dataset (TinyStories) — limited domain
- Fixed time budget training — not fully trained to convergence
- No RLHF or safety alignment
---
# Ethical Considerations
- This is a research artifact, not a production model
- The training data consists of synthetic stories (GPT-4 generated)
- No harmful content filtering was applied
- Intended for research and educational use only
---
# Citation
```bibtex
@misc{autoresearch_tinystories_depth8,
title={AutoResearch-tinystories-depth8},
author={Dustin Loring},
year={2026},
howpublished={\url{https://huggingface.co/quik-models/bumbling-wind-67}}
}}
```
---
# Version History
| Version | Date | Notes |
|----------|------|------|
| v1.0 | 2026-07-28 | Initial release |
---
# Acknowledgements
Built with the **AutoResearch** training framework.
Thanks to:
- Hugging Face
- PyTorch
- The creators of the TinyStories dataset
- The open-source AI research community
---
# License
This model is released under the **MIT License** unless otherwise specified.