File size: 3,285 Bytes
f4fd43c
1fd5981
 
f4fd43c
1fd5981
e2c11e0
1fd5981
 
 
 
 
 
1d17486
 
 
 
 
 
 
5d58380
1d17486
 
 
 
ab1c476
1d17486
1fd5981
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
---
language:
  - en
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
pipeline_tag: text-generation
tags:
  - gguf
  - transcript-cleaning
  - dictation
  - speech-to-text
  - llama.cpp
model-index:
  - name: flowbee-cut
    results:
      - task:
          type: text-generation
          name: Transcript cleaning
        dataset:
          name: Flowbee Cut eval battery (held-out, 35 cases)
          type: flowbee-cut-battery
        metrics:
          - name: Battery pass rate
            type: accuracy
            value: 1.0000
            verified: false
---

# Flowbee Cut — technical dictation cleaner

Fine-tune of [Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507)
that cleans raw speech-to-text transcripts for [Flowbee](https://github.com/auswm85/flowbee),
a local-first macOS dictation utility. It removes fillers and stutters,
resolves self-corrections, and writes technical speech in its correct form:

| spoken                                  | written                   |
| --------------------------------------- | ------------------------- |
| "rename it to camel case get user data" | Rename it to getUserData. |
| "run cargo test dash dash release"      | Run cargo test --release. |
| "open main dot rs"                      | Open main.rs.             |
| "we deploy behind engine x"             | We deploy behind nginx.   |

**This is not a chat model.** It was trained to do exactly one thing under
one system prompt, and it will clean — never answer — instruction-shaped
transcripts ("write a unit test for the auth module" comes back as cleaned
text, not a unit test).

## Usage contract

The model expects the exact Flowbee Cut system prompt it was trained with
(the `coder` prompt in `scripts/cut-eval/prompts.mjs` of the Flowbee repo),
with the raw transcript as the sole user message, `temperature 0`. Behavior
under other prompts is untested. Serve with llama.cpp:

```sh
llama-server -m flowbee-cut-<version>.Q4_K_M.gguf -ngl 99 -c 4096
```

## Training

- LoRA (r=16, attention projections, completion-only loss) on ~4,400
  synthetic pairs of messy spoken transcript → clean text: instruction-shaped
  technical dictation, CLI commands and flags, spoken identifiers and case
  directives, glossary-conditioned phonetic repairs, everyday dictation, and
  passthrough negatives. Adapter merged into the base weights, quantized to
  Q4_K_M.

## Files

- `flowbee-cut-<version>.Q4_K_M.gguf` — versioned releases (~2.5 GB).
- `latest.json` — machine-read manifest (version, file, sha256, eval score).
  The Flowbee app checks it on startup and downloads new releases, verifying
  the sha256 before the file touches a GGUF parser. Do not rename or delete
  these files by hand.

## Limitations

- **English only.** Training data is English; the base model is multilingual
  but this fine-tune's behavior on non-English transcripts is untested.
- Tuned for software-engineering vocabulary; exotic garbled jargon without a
  glossary hint is passed through verbatim by design (never deleted, never
  guessed).
- Trained on synthetic data seeded with real dictation failures; expect
  occasional misses on unusual phrasing (e.g. a garbled term directly
  adjacent to a modifier).

## License

Apache-2.0