auswm85 commited on
Commit
1fd5981
·
verified ·
1 Parent(s): a82cb8a

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +75 -5
README.md CHANGED
@@ -1,8 +1,78 @@
1
  ---
 
 
2
  license: apache-2.0
3
- base_model:
4
- - Qwen/Qwen3-4B-Instruct-2507
5
  pipeline_tag: text-generation
6
- language:
7
- - en
8
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - en
4
  license: apache-2.0
5
+ base_model: Qwen/Qwen3-4B-Instruct-2507
 
6
  pipeline_tag: text-generation
7
+ tags:
8
+ - gguf
9
+ - transcript-cleaning
10
+ - dictation
11
+ - speech-to-text
12
+ - llama.cpp
13
+ ---
14
+
15
+ # Flowbee Cut — technical dictation cleaner
16
+
17
+ Fine-tune of [Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507)
18
+ that cleans raw speech-to-text transcripts for [Flowbee](https://github.com/auswm85/flowbee),
19
+ a local-first macOS dictation utility. It removes fillers and stutters,
20
+ resolves self-corrections, and writes technical speech in its correct form:
21
+
22
+ | spoken | written |
23
+ | --------------------------------------- | ------------------------- |
24
+ | "rename it to camel case get user data" | Rename it to getUserData. |
25
+ | "run cargo test dash dash release" | Run cargo test --release. |
26
+ | "open main dot rs" | Open main.rs. |
27
+ | "we deploy behind engine x" | We deploy behind nginx. |
28
+
29
+ **This is not a chat model.** It was trained to do exactly one thing under
30
+ one system prompt, and it will clean — never answer — instruction-shaped
31
+ transcripts ("write a unit test for the auth module" comes back as cleaned
32
+ text, not a unit test).
33
+
34
+ ## Usage contract
35
+
36
+ The model expects the exact Flowbee Cut system prompt it was trained with
37
+ (the `coder` prompt in `scripts/cut-eval/prompts.mjs` of the Flowbee repo),
38
+ with the raw transcript as the sole user message, `temperature 0`. Behavior
39
+ under other prompts is untested. Serve with llama.cpp:
40
+
41
+ ```sh
42
+ llama-server -m flowbee-cut-<version>.Q4_K_M.gguf -ngl 99 -c 4096
43
+ ```
44
+
45
+ ## Training
46
+
47
+ - LoRA (r=16, attention projections, completion-only loss) on ~4,400
48
+ synthetic pairs of messy spoken transcript → clean text: instruction-shaped
49
+ technical dictation, CLI commands and flags, spoken identifiers and case
50
+ directives, glossary-conditioned phonetic repairs, everyday dictation, and
51
+ passthrough negatives. Adapter merged into the base weights, quantized to
52
+ Q4_K_M.
53
+ - Trained, evaluated, and published automatically by CI (GitHub Actions →
54
+ Modal GPU). A release is published only if it clears a held-out 22-case
55
+ evaluation battery; the current release scores **21/22**.
56
+
57
+ ## Files
58
+
59
+ - `flowbee-cut-<version>.Q4_K_M.gguf` — versioned releases (~2.5 GB).
60
+ - `latest.json` — machine-read manifest (version, file, sha256, eval score).
61
+ The Flowbee app checks it on startup and downloads new releases, verifying
62
+ the sha256 before the file touches a GGUF parser. Do not rename or delete
63
+ these files by hand.
64
+
65
+ ## Limitations
66
+
67
+ - **English only.** Training data is English; the base model is multilingual
68
+ but this fine-tune's behavior on non-English transcripts is untested.
69
+ - Tuned for software-engineering vocabulary; exotic garbled jargon without a
70
+ glossary hint is passed through verbatim by design (never deleted, never
71
+ guessed).
72
+ - Trained on synthetic data seeded with real dictation failures; expect
73
+ occasional misses on unusual phrasing (e.g. a garbled term directly
74
+ adjacent to a modifier).
75
+
76
+ ## License
77
+
78
+ Apache-2.0