Unseen1980's picture
instruct README.md
5e9572f verified
|
Raw
History Blame Contribute Delete
3.14 kB
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
  - daedalus
  - cpu-inference
  - gguf
  - q4_0

daedalus-150m

A 160.5M-parameter causal LM built for the best quality-per-token-per-second on CPU inference, exported to GGUF Q4_0 for llama.cpp.

What this model is trying to beat

Beat Pythia-160M, OPT-125M and GPT-neo-125M on quality; target MobileLLM-125M as a stretch; concede SmolLM2-135M on quality while beating it decisively on CPU decode.

This bar was fixed before any result landed. Numbers below are reported against it whether or not they clear it.

Architecture

exported as Lfm2ForCausalLM
parameters 160,488,960 (122,740,224 non-embedding)
blocks 18 (ccccAccAcAcAcAccAc -- c = gated short conv, A = GQA attention)
hidden size 768
SwiGLU inner dim 2048
heads 12 query / 4 KV, head_dim 64, QK-norm
RoPE theta 1,000,000
context 2048
tied embeddings True
tokenizer HuggingFaceTB/SmolLM2-135M, reused byte-identical, vocab 49,152

Training

  • run: post-sft
  • Muon lr 0.002 on 2D hidden matrices; AdamW lr 3e-05 on embeddings/head/norms
  • WSD schedule, linear decay to zero over the final 45% of the run

Evaluation

Not yet measured for this export.

Q4_0 quantization

  • fp16 perplexity 217.8803 vs Q4_0 263.4498
  • delta 20.915%
  • passes the <1.0% threshold: False
  • llama.cpp CPU decode at 64 threads, by context depth:
    • depth 0: 118.0 tok/s (+/- 3.8)
    • depth 512: 103.3 tok/s (+/- 3.4)
    • depth 2048 (the trained context): 97.5 tok/s (+/- 2.4)

Checkpoints and how to continue training

Checkpoints are pushed to the private Hub model repo Unseen1980/daedalus-checkpoints: weights-only bf16 rolling copies every ~2 h under rolling/<run>/weights.pt, plus a milestone with full Muon + AdamW optimizer state at the WSD decay-start step on its own revision.

No milestone record was found beside this checkpoint, so no branch point is published for it.

Deviations from the blueprint

Each was costed and approved rather than silently dropped; see DAEDALUS-BLUEPRINT-v6.md and issue #4.

  • No distillation from SmolLM2-1.7B during decay. 288 GB of top-16 logits does not fit the disk and the online-teacher variant cost ~$29 of a $94.66 budget; its own evidence was only "+1-3 points plausible".
  • Corpus stops at ~14.2B tokens, not 45B. Training repeats a balanced corpus rather than seeing 45B unique tokens; at this scale repetition up to ~4 epochs costs little against fresh tokens, and mixture balance mattered more than raw size.
  • Document-aligned packing not implemented -- sequences may cross document boundaries.
  • NoPE skipped -- it breaks GGUF export.
  • Single seed for the hero run, so no seed-sigma is reported.
  • everyday-conversations contributes ~0.00% of pretraining instead of its 2% share (the whole dataset is 0.4M tokens, which the 4-epoch cap reduces to nothing); dialogue enters at the post SFT stage instead.