aetherstory / README.md
wincode's picture
Upload README.md with huggingface_hub
2741fc9 verified
|
Raw
History Blame Contribute Delete
2.08 kB
metadata
language:
  - en
license: mit
library_name: custom
tags:
  - from-scratch
  - storytelling
  - creative-writing
  - cpu-trained
  - transformer
  - pytorch
base_model: []
pipeline_tag: text-generation

AetherStory

A from-scratch, CPU-trained storyteller transformer. ~863,492 parameters.

AetherStory is a tiny decoder-only transformer (GPT-style) trained entirely on CPU on the procedurally generated AetherStory dataset. It writes short fantasy fables given an opening prompt.

This is not a fine-tune of a larger model and not a wrapper around transformers — every layer is implemented by hand in plain PyTorch.

Architecture

 token + position embeddings
            |
   +----------------+
   | Transformer x4 |
   |  causal MHA     |
   |  GELU FFN       |
   +----------------+
            |
        LayerNorm
            |
   tied output head
hyperparameter value
vocab size 10000
d_model 128
layers 4
heads 4
ffn dim 512
max seq len 64
parameters 863,492
tied embeddings True

Training

Trained with AdamW (lr 3e-4, cosine schedule, warmup 200) for 4 epochs on a 4-core CPU. Best validation loss: 0.3103, trained in unknown (recovered) on CPU.

loss curve

Usage

# files needed next to this script:
#   model.safetensors, config.json, tokenizer.json, modeling_aetherstory.py
from modeling_aetherstory import StoryTeller

teller = StoryTeller.from_dir(".")
print(teller("In the Glasslands there lived", max_tokens=100, temperature=0.9))

Limitations

A ~2M-parameter model trained on synthetic fables will not produce literature. It will produce charming, sometimes incoherent, fairy-tale-flavoured text — which is the point. It is a demonstration that a small, fully custom model can be trained, evaluated, and shipped end-to-end on commodity hardware.

License

MIT.