mumble-cleanup / DEVELOPMENT.md
adikuma's picture
initial upload: cleanup code and 688-pair seed dataset
fd0b01f verified
|
Raw
History Blame Contribute Delete
5.68 kB

cleanup

Status: implementation. Run make smoke first to validate wiring on CPU. Run make all RUN_ID=r1 on a Vast.ai GPU for the real training pass. Run make benchmark locally on a laptop for the latency truth.

Fine-tune Qwen/Qwen2.5-0.5B-Instruct with LoRA to clean Mumble dictation transcripts. Remove fillers, repeats, and false starts. Fix punctuation and capitalization. Never paraphrase, summarize, or add content. Export to ONNX (fp32 + int8) for the Mumble Rust app's ort runtime.

See:

Pipeline

make sync                  install deps via uv
make data                  01: split the seed jsonl into train/val/test under data/pairs/
make train RUN_ID=r1       02: lora sft on qwen2.5-0.5b-instruct
make evaluate RUN_ID=r1    03: raw vs base vs fine-tune on the held-out test, plus adversarial check
make export RUN_ID=r1      04: merge lora and export fp32 + int8 onnx
make benchmark RUN_ID=r1   05: cpu latency, LOCAL only
make pack RUN_ID=r1        06: tar the run, print sha256 and the vastai copy line
make push RUN_ID=r1        publish to the hf hub
make all RUN_ID=r1         pipeline end to end on the gpu (no benchmark, no push)
make smoke                 tiny cpu run, no gpu needed

Data

Source of truth: data/seed/synthetic_pairs.jsonl β€” 688 hand-curated (raw, clean) pairs generated by a multi-agent workflow across 8 categories of dictation:

  • casual_messages β€” Slack / texts
  • professional_emails β€” work emails
  • meeting_notes β€” voice memos during/after meetings
  • technical_dictation β€” code comments, API notes, design sketches
  • todo_lists β€” action items, shopping, plans
  • long_form_thoughts β€” paragraph-length thinking out loud
  • questions_and_asks β€” questions, requests, asks
  • mixed_content β€” names, numbers, dates, times, addresses

Each pair is { raw: <Parakeet-shaped lowercase no-punct disfluent input>, clean: <proper English output> }. The clean side is faithful by construction: every content word in clean exists in raw (modulo homophone fixes). The model is mechanically incapable of learning to invent content because the data never rewards it.

scripts/01_download.py splits the seed stratified-by-category into train/val/test (85/10/5) and writes data/pairs/{train,val,test}.json plus data/pairs/meta.json.

The 9 synthetic disfluency operators in src/cleanup/data/inject.py are kept for an optional v1.1 augmentation pass that expands the seed by 3-5x by stochastically corrupting clean sides. Off by default in v1.

Layout

configs/
  data.yaml                seed path + split ratios
  inject.yaml              optional operator probabilities (v1.1)
  train.yaml               lora + adamw hyperparameters
src/cleanup/
  config.py                yaml -> typed dataclasses
  prompts.py               frozen system prompt
  infer.py                 model -> cleaned string (one-shot helper)
  data/
    download.py            load seed jsonl, split stratified, write data/pairs/
    inject.py              9 disfluency operators (v1.1 augmentation, off by default)
    tokenize.py            chatml format + the response_template for completion-only loss
  train/
    model.py               load qwen + apply lora + autodetect precision
    trainer.py             trl sftrainer wiring with completion-only collator
  eval/
    metrics.py             disfluency removal, punct f1, faithfulness, length, pass rate
    latency.py             cpu p50/p95/p99 sweep + realistic mix
    report.py              read jsons -> charts + markdown
  export/
    merge.py               peft merge_and_unload
    to_onnx.py             optimum onnx export
    quantize.py            ort int8 dynamic quantization
  pack/
    ship.py                tar + sha256 + vastai copy line
scripts/
  01_download.py           split seed jsonl into train/val/test
  02_train.py              lora sft on qwen2.5-0.5b-instruct
  03_evaluate.py           raw vs base vs fine-tune + adversarial questions
  04_export.py             merge + onnx + int8
  05_benchmark.py          cpu latency (LOCAL only)
  06_pack_and_ship.py      tar + sha256 + vastai copy line
  push_to_hub.py           hf hub publish (private by default)
docs/
  model_card.md            hf hub readme (with frontmatter)
  model_report.md          human-readable report (numbers filled after a run)
  HANDBOOK.md              deep specs
data/
  seed/synthetic_pairs.jsonl   the 612-pair source of truth
  seed/synthetic_pairs.csv     same data, human-inspectable
  seed/batches/                per-category breakdowns
  pairs/{train,val,test}.json  created by scripts/01_download.py
runs/<run-id>/
  model/                       peft adapter + tokenizer
  merged/                      transformers checkpoint after lora merge
  onnx/model.onnx              fp32 merged onnx
  onnx/int8/model.onnx         int8 quantized onnx
  eval.json                    quality metrics (raw, base, fine_tuned, adversarial)
  latency_benchmark.json       cpu latency (only after local benchmark)
  training_history.json        loss curves
dist/                          packed tarballs ready to scp