metadata
license: apache-2.0
datasets:
- mookiezi/Discord-Dialogues
- openbmb/Ultra-FineWeb-L1
base_model: ProCreations/booper-pretrain
language:
- en
tags:
- babble
- booper
- conversational
- discord
library_name: pytorch
pipeline_tag: text-generation
This model was asked to be published under my account, not the creators. The compute came from https://huggingface.co/posts/ProCreations/855858308074329
Booper chat
A small chat model: the babble / booper transformer, pretrained on English web text, then supervised-fine-tuned on Discord conversations so it replies instead of only continuing prose.
- Base: ProCreations/booper-pretrain (34.1M params, Ultra-FineWeb-L1)
- Post-train data: mookiezi/Discord-Dialogues (ChatML two-author Discord threads)
- This is still a tiny model. It will sound like Discord, and it will be clumsy.
How to prompt it
Training used babble's pair layout, with loss only on the assistant side:
<bos> {user message} <sep> {reply} <eos>
At inference, feed <bos> your message <sep> and sample until <eos>. Multi-turn: put earlier turns in the prompt (newlines), last user message at the end, then <sep>.
Training
| Parameters | 34,096,128 (8 layers, 512 wide, 8 heads, context 1024) |
| Tokenizer | same byte-level BPE as the pretrain (16,384 tokens) |
| Objective | SFT, assistant tokens only |
| Assistant tokens | 150,001,494 |
| Steps | 36,017 |
| Hardware | 1Γ NVIDIA H200 (Hugging Face Jobs) |
| Wall clock | |
| Val loss | 3.05 β 2.590 |
Job: ProCreations/6a8a1cc87c5c7dd37923553d
End-of-run samples
Temperature 0.7, top-k 40:
heyβHii hruwhat's upβI just got on matelolβI think they both are too smallcan you help meβwhat level is that
Files
latest.ptβ checkpoint (model,config, optimizer,stage: posttrain-discord)tokenizer.jsonβ BPE merges (babble.subword.BPETokenizer.from_json)loss.jsonlβ val loss, throughput, and samples per checkpointrun_meta.jsonβ dataset / LR / batch used for this run
Source architecture and tokenizer scheme: kowo-co/babble.