Tatarby models

Every model that Актәпи (tatarby), a Tatar voice assistant, runs locally, in one place, so the app's ./start.sh can fetch them all with one command. Each file is an unmodified copy of its upstream release, apart from voice-reference.pt, which the project builds itself. The SHA-256 of each file matches the upstream file and is pinned in the project's setup scripts.

File What it does Upstream License
asr/ggml-tt-small-q4_0.bin (139 MB) Tatar speech recognition (whisper.cpp) JackGauter/whisper-small-tt-ggml @ 95644ef, quantized from yasalma/whisper-finetuned-tt-asr Apache-2.0
asr/ggml-large-v3-turbo-q5_0.bin (547 MB) Russian speech recognition (whisper.cpp) ggerganov/whisper.cpp @ 5359861 (OpenAI Whisper large-v3-turbo) MIT
tts/mms-tts-tat/ (139 MB) Tatar speech synthesis (VITS) facebook/mms-tts-tat @ a7c206e CC BY-NC 4.0
tts/silero-v5-5-ru.pt (139 MB) Russian speech synthesis Silero v5_5_ru CC BY-NC-SA 4.0
tts/knn-vc/WavLM-Large.pt (1.2 GB) Speech features for voice conversion bshall/knn-vc release v0.1 (WavLM, Microsoft) MIT
tts/knn-vc/prematch_g_02500000.pt (63 MB) Vocoder for voice conversion (HiFi-GAN, prematched) bshall/knn-vc release v0.1 MIT
tts/knn-vc/voice-reference.pt (36 MB) WavLM features of the Tatar voice reading 170 Tatar and Russian sentences: Russian speech gets converted into this voice built by tatarby-tts/scripts/build_voice_reference.py from mms-tts-tat output CC BY-NC 4.0 (derived from mms-tts-tat)

Non-commercial use only. mms-tts-tat, Silero and voice-reference.pt are under non-commercial licenses, so the full set may only be used non-commercially. See LICENSE.md.

Download

The app's ./start.sh downloads missing files on first run. By hand:

hf download JackGauter/tatarby-models --local-dir tatarby-models

Citations

  • Pratap et al., Scaling Speech Technology to 1,000+ Languages (MMS), 2023.
  • Baas, van Niekerk, Kamper, Voice Conversion With Just Nearest Neighbors (kNN-VC), Interspeech 2023.
  • Chen et al., WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, 2022.
  • Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper), 2022.
  • Silero Team, Silero Models, https://github.com/snakers4/silero-models.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support