Tatarby models
Every model that Актәпи (tatarby), a Tatar voice assistant, runs locally, in one place, so the app's ./start.sh can fetch them all with one command. Each file is an unmodified copy of its upstream release, apart from voice-reference.pt, which the project builds itself. The SHA-256 of each file matches the upstream file and is pinned in the project's setup scripts.
| File | What it does | Upstream | License |
|---|---|---|---|
asr/ggml-tt-small-q4_0.bin (139 MB) |
Tatar speech recognition (whisper.cpp) | JackGauter/whisper-small-tt-ggml @ 95644ef, quantized from yasalma/whisper-finetuned-tt-asr |
Apache-2.0 |
asr/ggml-large-v3-turbo-q5_0.bin (547 MB) |
Russian speech recognition (whisper.cpp) | ggerganov/whisper.cpp @ 5359861 (OpenAI Whisper large-v3-turbo) |
MIT |
tts/mms-tts-tat/ (139 MB) |
Tatar speech synthesis (VITS) | facebook/mms-tts-tat @ a7c206e |
CC BY-NC 4.0 |
tts/silero-v5-5-ru.pt (139 MB) |
Russian speech synthesis | Silero v5_5_ru |
CC BY-NC-SA 4.0 |
tts/knn-vc/WavLM-Large.pt (1.2 GB) |
Speech features for voice conversion | bshall/knn-vc release v0.1 (WavLM, Microsoft) | MIT |
tts/knn-vc/prematch_g_02500000.pt (63 MB) |
Vocoder for voice conversion (HiFi-GAN, prematched) | bshall/knn-vc release v0.1 | MIT |
tts/knn-vc/voice-reference.pt (36 MB) |
WavLM features of the Tatar voice reading 170 Tatar and Russian sentences: Russian speech gets converted into this voice | built by tatarby-tts/scripts/build_voice_reference.py from mms-tts-tat output |
CC BY-NC 4.0 (derived from mms-tts-tat) |
Non-commercial use only. mms-tts-tat, Silero and voice-reference.pt are under non-commercial licenses, so the full set may only be used non-commercially. See LICENSE.md.
Download
The app's ./start.sh downloads missing files on first run. By hand:
hf download JackGauter/tatarby-models --local-dir tatarby-models
Citations
- Pratap et al., Scaling Speech Technology to 1,000+ Languages (MMS), 2023.
- Baas, van Niekerk, Kamper, Voice Conversion With Just Nearest Neighbors (kNN-VC), Interspeech 2023.
- Chen et al., WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, 2022.
- Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper), 2022.
- Silero Team, Silero Models, https://github.com/snakers4/silero-models.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support