booper-chat / README.md
ProCreations's picture
Fill in Discord post-train results and chat prompt format
178cfb3 verified
|
Raw
History Blame Contribute Delete
2.46 kB
metadata
license: apache-2.0
datasets:
  - mookiezi/Discord-Dialogues
  - openbmb/Ultra-FineWeb-L1
base_model: ProCreations/booper-pretrain
language:
  - en
tags:
  - babble
  - booper
  - conversational
  - discord
library_name: pytorch
pipeline_tag: text-generation

This model was asked to be published under my account, not the creators. The compute came from https://huggingface.co/posts/ProCreations/855858308074329

Booper chat

A small chat model: the babble / booper transformer, pretrained on English web text, then supervised-fine-tuned on Discord conversations so it replies instead of only continuing prose.

How to prompt it

Training used babble's pair layout, with loss only on the assistant side:

<bos> {user message} <sep> {reply} <eos>

At inference, feed <bos> your message <sep> and sample until <eos>. Multi-turn: put earlier turns in the prompt (newlines), last user message at the end, then <sep>.

Training

Parameters 34,096,128 (8 layers, 512 wide, 8 heads, context 1024)
Tokenizer same byte-level BPE as the pretrain (16,384 tokens)
Objective SFT, assistant tokens only
Assistant tokens 150,001,494
Steps 36,017
Hardware 1Γ— NVIDIA H200 (Hugging Face Jobs)
Wall clock 73 minutes (72 min of training at ~34.6k target-tok/s)
Val loss 3.05 β†’ 2.590

Job: ProCreations/6a8a1cc87c5c7dd37923553d

End-of-run samples

Temperature 0.7, top-k 40:

  • hey β†’ Hii hru
  • what's up β†’ I just got on mate
  • lol β†’ I think they both are too small
  • can you help me β†’ what level is that

Files

  • latest.pt β€” checkpoint (model, config, optimizer, stage: posttrain-discord)
  • tokenizer.json β€” BPE merges (babble.subword.BPETokenizer.from_json)
  • loss.jsonl β€” val loss, throughput, and samples per checkpoint
  • run_meta.json β€” dataset / LR / batch used for this run

Source architecture and tokenizer scheme: kowo-co/babble.