--- license: apache-2.0 datasets: - mookiezi/Discord-Dialogues - openbmb/Ultra-FineWeb-L1 base_model: ProCreations/booper-pretrain language: - en tags: - babble - booper - conversational - discord library_name: pytorch pipeline_tag: text-generation --- This model was asked to be published under my account, not the creators. The compute came from https://huggingface.co/posts/ProCreations/855858308074329 # Booper chat A small chat model: the [babble / booper](https://github.com/kowo-co/babble) transformer, pretrained on English web text, then supervised-fine-tuned on Discord conversations so it replies instead of only continuing prose. - Base: [ProCreations/booper-pretrain](https://huggingface.co/ProCreations/booper-pretrain) (34.1M params, Ultra-FineWeb-L1) - Post-train data: [mookiezi/Discord-Dialogues](https://huggingface.co/datasets/mookiezi/Discord-Dialogues) (ChatML two-author Discord threads) - This is still a tiny model. It will sound like Discord, and it will be clumsy. ## How to prompt it Training used babble's pair layout, with loss only on the assistant side: ``` {user message} {reply} ``` At inference, feed ` your message ` and sample until ``. Multi-turn: put earlier turns in the prompt (newlines), last user message at the end, then ``. ## Training | | | |---|---| | Parameters | 34,096,128 (8 layers, 512 wide, 8 heads, context 1024) | | Tokenizer | same byte-level BPE as the pretrain (16,384 tokens) | | Objective | SFT, assistant tokens only | | Assistant tokens | 150,001,494 | | Steps | 36,017 | | Hardware | 1× NVIDIA H200 (Hugging Face Jobs) | | Wall clock | ~73 minutes (~72 min of training at ~34.6k target-tok/s) | | Val loss | 3.05 → **2.590** | Job: [ProCreations/6a8a1cc87c5c7dd37923553d](https://huggingface.co/jobs/ProCreations/6a8a1cc87c5c7dd37923553d) ## End-of-run samples Temperature 0.7, top-k 40: - `hey` → `Hii hru` - `what's up` → `I just got on mate` - `lol` → `I think they both are too small` - `can you help me` → `what level is that` ## Files - `latest.pt` — checkpoint (`model`, `config`, optimizer, `stage: posttrain-discord`) - `tokenizer.json` — BPE merges (`babble.subword.BPETokenizer.from_json`) - `loss.jsonl` — val loss, throughput, and samples per checkpoint - `run_meta.json` — dataset / LR / batch used for this run Source architecture and tokenizer scheme: [kowo-co/babble](https://github.com/kowo-co/babble).