stillerman's picture
publish disfluency tagger
a884191 verified
|
Raw
History Blame Contribute Delete
1.11 kB
---
license: apache-2.0
pipeline_tag: token-classification
tags:
- disfluency-detection
- transformers.js
- onnx
---
# stillerman/fdt-disfluency-medium-41m
Disfluency **deletion tagger** for live speech transcripts: tags every
whitespace word `KEEP` / `DELETE` / `KEEP_STRIP_COMMA` / `KEEP_CAPITALIZE`,
then a ~15-line reconstruction turns tags into cleaned text. Deletion-only by
construction — it cannot rephrase, hallucinate, or alter names and numbers.
- Architecture: BERT L8/H512 (41M params), v1 data mix
- Val metrics: exact-match 0.9455, DELETE-F1 0.9887
- Training data: synthetic disfluency injection over conversational corpora —
see [stillerman/fdt-disfluency-synthetic](https://huggingface.co/datasets/stillerman/fdt-disfluency-synthetic)
- `onnx/model_quantized.onnx` (int8) is ready for transformers.js
(`device: "webgpu"`, `dtype: "q8"`); runs at ~10–50 ms per utterance
in-browser.
- ⚠️ Trained partly on DailyDialog (CC BY-NC-SA): treat as research
artifact, not for commercial deployment as-is.
Trained on a DGX Spark as part of the FluencyAI digital-twin project.