neat-elevator-1 / README.md
quik-models's picture
Upload autoresearch experiment (val_bpb=2.415760)
7af3a9f verified
|
Raw
History Blame Contribute Delete
8.84 kB
---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- causal-lm
- pytorch
- small-language-model
- autoresearch
- from-scratch
datasets:
- multivision
---
# AutoResearch-multivision-depth8
![AutoResearch Cover](https://raw.githubusercontent.com/nullbotai-droid/turbo-diffusion-local/refs/heads/main/cover2.png)
**AutoResearch-multivision-depth8** is a **77.6M parameter** decoder-only Transformer trained **from scratch** on **multivision**.
This model is part of the **AutoResearch** project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
---
# Overview
This is a 8-layer decoder-only Transformer trained on the multivision dataset for 1.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 2.415760 (perplexity: 5.3360) on the held-out validation set.
---
# References
## Papers
- NanoGPT / NanoChat architecture patterns
## Datasets
- Training: multivision
- Tokenizer: multivision
## Related Projects
- [karpathy/autoresearch](https://github.com/karpathy/autoresearch)
## WANDB Run
- [https://wandb.ai/dustinsdelivery-delivery-ai/vision-experiment/runs/0meyvt63](https://wandb.ai/dustinsdelivery-delivery-ai/vision-experiment/runs/0meyvt63)
---
# Highlights
- Trained **from scratch**
- **77.6M parameters**
- Trained on **111.7M tokens** (213 steps)
- 8-layer decoder-only Transformer with sliding window attention
- RoPE positional encoding, RMSNorm, ReLU² activation
- MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
- Hugging Face Transformers compatible
---
# Model Architecture
| Property | Value |
|-----------|------:|
| Architecture | Decoder-only Transformer |
| Parameters | **77,575,312** (77.6M) |
| Layers | 8 |
| Hidden Size | 512 |
| Attention Heads | 4 |
| KV Heads | 4 |
| Head Dimension | 128 |
| Feed Forward Size | 2048 |
| Context Length | 2048 |
| Vocabulary Size | 16,384 |
| Positional Encoding | RoPE |
| Activation | ReLU² |
| Normalization | RMSNorm |
| Window Pattern | SSSL |
| Weight Tying | No |
---
# Training
This model was trained **from scratch** for **1.0 hours** (3606s) of wall-clock training time.
## Training Configuration
| Setting | Value |
|---------|------:|
| Optimizer | MuonAdamW (Muon + AdamW) |
| Precision | torch.bfloat16 |
| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
| Weight Decay | 0.2 |
| Batch Size | 4 × 2048 = 8,192 tokens/step |
| Gradient Accumulation | 64 steps |
| Total Batch Size | 524,288 tokens |
| Context Length | 2048 |
| Vocabulary | 16,384 tokens (BPE) |
| LR Scheduler | Linear warmdown (50%) |
| Activation Checkpointing | Enabled |
## Hardware
- GPU: NVIDIA GeForce RTX 4060 Ti
- VRAM: 16.0 GB
- Peak VRAM Used: 3.7 GB
- MFU: 9.24%
- Framework: PyTorch 2.9.1+cu128
---
# Dataset
- **Name:** multivision
- **Language:** English
### Preprocessing
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
---
# Intended Use
This model is intended for:
- Educational purposes and research
- Text generation experiments
- Studying small language model training dynamics
Not recommended for:
- Production use or safety-critical applications
- Tasks requiring factual accuracy
---
# Evaluation
## Results
| Metric | Score |
|---------|------:|
| Validation BPB | 2.415760 |
| Perplexity | 5.3360 |
| Peak VRAM | 3.7 GB |
| MFU | 9.24% |
---
# Example Generations
## Example 1
### Prompt
```text
Once upon a time,
```
### Generation
```text
Once upon a time, and the overall composition highlights both functionality and aesthetic appeal.
```
## Example 2
### Prompt
```text
A lonely dragon
```
### Generation
```text
A lonely dragon dragon dragon dragon lion dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon dragon
```
## Example 3
### Prompt
```text
The opposite of boy is
```
### Generation
```text
The opposite of boy is a child wearing a bright pink and white striped shirt. The background shows a blurred urban setting with buildings and trees, suggesting a city environment. The overall scene conveys a sense of a community gathering or a simple event.
```
## Example 4
### Prompt
```text
The opposite of queen is
```
### Generation
```text
The opposite of queen is a vibrant red color with white accents. The background features a clear sky and a serene landscape, suggesting a peaceful day at the beach. The overall composition conveys a sense of tranquility and appreciation for nature's beauty.
```
## Example 5
### Prompt
```text
My name is
```
### Generation
```text
My name is prominently displayed in the foreground. The background
```
## Example 6
### Prompt
```text
2 + 2 is
```
### Generation
```text
2 + 2 is 2. 4 17 001 3 8. 4. 5 2- 5 2 900 18 19 2 25 3 8 00 3 2 1 00 1
```
# Usage
```python
import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer
# Load config
with open('config.json', 'r') as f:
config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()
# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
tokenizer = pickle.load(f)
# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
for _ in range(50):
logits = model(x)
probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
input_ids.append(next_token.item())
x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))
```
---
# Repository Structure
```text
model.pt # Model weights
config.json # Model architecture config
dataset.txt # Dataset name used for training
token_bytes.pt # Token byte mappings
tokenizer.pkl # Trained BPE tokenizer
tokenizer_config.json # Tokenizer configuration
training_metrics.json # Training metrics
README.md # This file
```
---
# Limitations
- Small model size limits language understanding and coherence
- Trained on a single dataset (TinyStories) — limited domain
- Fixed time budget training — not fully trained to convergence
- No RLHF or safety alignment
---
# Ethical Considerations
- This is a research artifact, not a production model
- The training data consists of synthetic stories (GPT-4 generated)
- No harmful content filtering was applied
- Intended for research and educational use only
---
# Citation
```bibtex
@misc{autoresearch_multivision_depth8,
title={AutoResearch-multivision-depth8},
author={Dustin Loring},
year={2026},
howpublished={\url{https://huggingface.co/quik-models/neat-elevator-1}}
}}
```
---
# Version History
| Version | Date | Notes |
|----------|------|------|
| v1.0 | 2026-07-31 | Initial release |
---
# Acknowledgements
Built with the **AutoResearch** training framework.
Thanks to:
- Hugging Face
- PyTorch
- The creators of the TinyStories dataset
- The open-source AI research community
---
# License
This model is released under the **MIT License** unless otherwise specified.