This model was asked to be published under my account, not the creators. The compute came from https://huggingface.co/posts/ProCreations/855858308074329

Booper chat

A small chat model: the babble / booper transformer, pretrained on English web text, then supervised-fine-tuned on Discord conversations so it replies instead of only continuing prose.

How to prompt it

Training used babble's pair layout, with loss only on the assistant side:

<bos> {user message} <sep> {reply} <eos>

At inference, feed <bos> your message <sep> and sample until <eos>. Multi-turn: put earlier turns in the prompt (newlines), last user message at the end, then <sep>.

Training

Parameters 34,096,128 (8 layers, 512 wide, 8 heads, context 1024)
Tokenizer same byte-level BPE as the pretrain (16,384 tokens)
Objective SFT, assistant tokens only
Assistant tokens 150,001,494
Steps 36,017
Hardware 1Γ— NVIDIA H200 (Hugging Face Jobs)
Wall clock 73 minutes (72 min of training at ~34.6k target-tok/s)
Val loss 3.05 β†’ 2.590

Job: ProCreations/6a8a1cc87c5c7dd37923553d

End-of-run samples

Temperature 0.7, top-k 40:

  • hey β†’ Hii hru
  • what's up β†’ I just got on mate
  • lol β†’ I think they both are too small
  • can you help me β†’ what level is that

Files

  • latest.pt β€” checkpoint (model, config, optimizer, stage: posttrain-discord)
  • tokenizer.json β€” BPE merges (babble.subword.BPETokenizer.from_json)
  • loss.jsonl β€” val loss, throughput, and samples per checkpoint
  • run_meta.json β€” dataset / LR / batch used for this run

Source architecture and tokenizer scheme: kowo-co/babble.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ProCreations/booper-chat

Finetuned
(1)
this model

Datasets used to train ProCreations/booper-chat