File size: 6,191 Bytes
a168966 9caed08 a168966 4b7b9c2 2b9a24a a168966 401a1a6 a168966 db1992d a168966 9caed08 a168966 9caed08 a168966 9caed08 db1992d a168966 db1992d a168966 401a1a6 a168966 db1992d a168966 9caed08 a168966 9caed08 a168966 db1992d a168966 db1992d a168966 db1992d a168966 9caed08 a168966 db1992d a168966 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
license: cc-by-nc-4.0
tags:
- speech-restoration
- speech-enhancement
- telephony
- call-center
- w2v-bert
- lora
- sidon
- dac
base_model: facebook/w2v-bert-2.0
pipeline_tag: audio-to-audio
---
# CallEnhancer β Call-Centre / Telephony Speech Restoration
Restore Multi-lingual **narrowband, codec'd, noisy call-centre / telephony speech** (e.g. 8 kHz G.711/GSM phone
audio) to **clean 48 kHz**.
## Benchmark β CER
Character Error Rate (Whisper large-v3, **lower is better**) on a held-out **private call-centre
test set**. CallEnhancer restores telephony-degraded speech so an ASR reads it far more accurately
than the raw input or general-purpose speech enhancers.

| model | CER % β |
|---|---:|
| degraded (original) | 71.48 |
| voicefixer | 58.27 |
| resemble-enhance | 54.96 |
| sidon-v0.1 | 41.06 |
| **CallEnhancer-small** (this repo, open source) | **38.91** |
| **CallEnhancer-base** | **19.01** |
This repository hosts **CallEnhancer-small** (open source); **CallEnhancer-base** is the larger model.
## Quick start β infer from the HF checkpoint
```bash
pip install torch torchaudio "transformers>=4.56" "descript-audio-codec>=1.0.0" soundfile "huggingface_hub[cli]"
# pull the CLI + the two slim checkpoints from the Hub
hf auth login # private repo: log in first (or export HF_TOKEN=hf_...)
hf download Scicom-intl/CallEnhancer \
infer_callcentre.py fe_callcentre/fe_adapter_full.pt \
decoder_callcentre/decoder_only.pt --local-dir CallEnhancer
cd CallEnhancer && python infer_callcentre.py \
--input your_call.wav --out-dir out \
--fe-adapter fe_callcentre/fe_adapter_full.pt \
--decoder decoder_callcentre/decoder_only.pt \
--chunk 0 --device cuda # --chunk 0 = NO chunking (default single pass); --device cpu if no GPU
# -> out/your_call_restored48k.wav (clean 48 kHz) + out/your_call_orig48k.wav (A/B)
```
Prefer Python (load weights from the Hub with `hf_hub_download`)? See **[Python](#python-pull-weights-from-the-hub)** below.
> **Status:** CallEnhancer-small is fully trained and restores real 8 kHz call-centre audio well.
## Files
Use the **current-run** checkpoints under `fe_callcentre/` and `decoder_callcentre/`:
| path | role | size |
|---|---|---|
| `fe_callcentre/fe_adapter_full.pt` | **FE adapter (inference)** β 144 tensors: 96 LoRA + 48 trained `output_dense` biases | ~63 MB |
| `decoder_callcentre/decoder_only.pt` | **decoder (inference)** β 188M DAC decoder | ~0.75 GB |
| `fe_callcentre/last.pt`, `decoder_callcentre/last.pt` | raw checkpoints (resume training) | ~2.5 / 2.8 GB |
| `infer_callcentre.py` | inference CLI (below) | β |
For inference you only need the two slim files + `infer_callcentre.py`. *(Root-level
`fe_adapter_full.pt` / `decoder_only.pt` are from an earlier run and are superseded.)*
## End-to-end example (straight from HuggingFace)
```bash
pip install torch torchaudio "transformers>=4.56" "descript-audio-codec>=1.0.0" soundfile "huggingface_hub[cli]"
# pull the CLI + the two slim checkpoints, straight from this repo
hf auth login # private repo: log in first (or export HF_TOKEN=hf_...)
hf download Scicom-intl/CallEnhancer \
infer_callcentre.py \
fe_callcentre/fe_adapter_full.pt \
decoder_callcentre/decoder_only.pt \
--local-dir CallEnhancer
cd CallEnhancer
# restore your audio end-to-end
python infer_callcentre.py \
--input your_call.wav \
--out-dir out \
--fe-adapter fe_callcentre/fe_adapter_full.pt \
--decoder decoder_callcentre/decoder_only.pt \
--chunk 0 --device cuda # --chunk 0 = NO chunking (single straight pass, default); --device cpu if no GPU
```
Outputs:
- `out/your_call_restored48k.wav` β the restored **clean 48 kHz** speech.
- `out/your_call_orig48k.wav` β the input, naively upsampled to 48 kHz (no model), for an A/B listen.
`--input` accepts a **file or a directory** (`.wav/.flac/.mp3/.ogg/.opus/.m4a`). Stereo (e.g.
agent/customer on separate channels) is restored per channel and recombined.
### Python (pull weights from the Hub)
```python
import numpy as np, soundfile as sf, torch, torchaudio
from huggingface_hub import hf_hub_download
from transformers import AutoFeatureExtractor, Wav2Vec2BertModel
import dac
REPO, SSL, FE_SR, SR_OUT = "Scicom-intl/CallEnhancer", "facebook/w2v-bert-2.0", 16000, 48000
dev = torch.device("cuda" if torch.cuda.is_available() else "cpu")
ck = torch.load(hf_hub_download(REPO, "fe_callcentre/fe_adapter_full.pt"), map_location="cpu")
ad, scale = ck["adapter"], ck["lora_alpha"] / ck["r"]
fe = Wav2Vec2BertModel.from_pretrained(SSL, num_hidden_layers=ck.get("layers", 24), layerdrop=0.0)
sd = fe.state_dict() # merge LoRA -> base (no peft needed)
for p in sorted({k[:-len(".lora_A.default.weight")] for k in ad if k.endswith(".lora_A.default.weight")}):
sd[p+".weight"] = sd[p+".weight"].float() + scale * (ad[p+".lora_B.default.weight"].float() @ ad[p+".lora_A.default.weight"].float())
if p+".base_layer.bias" in ad: sd[p+".bias"] = ad[p+".base_layer.bias"].to(sd[p+".bias"].dtype)
fe.load_state_dict(sd); fe.to(dev).eval()
dck = torch.load(hf_hub_download(REPO, "decoder_callcentre/decoder_only.pt"), map_location="cpu")
dec = dac.model.dac.Decoder(input_channel=1024, channels=dck.get("dec_channels", 3072), rates=[8,5,4,3,2])
dec.load_state_dict(dck["decoder"]); dec.to(dev).eval()
proc = AutoFeatureExtractor.from_pretrained(SSL)
@torch.no_grad()
def restore(path, out="restored48k.wav"): # single straight pass
x, sr = sf.read(path, dtype="float32"); x = x.mean(1) if x.ndim > 1 else x
if sr != FE_SR: x = torchaudio.functional.resample(torch.from_numpy(x)[None], sr, FE_SR)[0].numpy()
x = x / (np.abs(x).max() + 1e-9) * 0.95
feats = {k: v.to(dev) for k, v in proc(x, sampling_rate=FE_SR, return_tensors="pt").items()}
y = dec(fe(**feats).last_hidden_state.transpose(1, 2)).squeeze().float().cpu().numpy()
sf.write(out, y / (np.abs(y).max() + 1e-9) * 0.97, SR_OUT); print("wrote", out)
restore("your_call.wav") # <-- your own telephony/call-centre audio
```
|