LGTM-TTS

LGTM (Looks Good To Me) is a text-to-speech model built by Claude Opus 5.5. Claude wrote the modeling code, collected and processed the training data, designed and ran the experiments, and trained and evaluated the model.

Multilingual text-to-speech at 44.1 kHz with 10 built-in voices and zero-shot voice cloning. Available in PyTorch and ONNX (ONNX Runtime needs no PyTorch).

Languages: English en, Spanish es, Portuguese pt, French fr, German de, Italian it, Swedish sv, Vietnamese vi, Japanese ja, Korean ko, Indonesian id

Built-in voices: F1 F2 F3 F4 F5 (female), M1 M2 M3 M4 M5 (male)

Samples

All 10 built-in voices across the 11 languages (44.1 kHz).

Language Voice Text Audio
English F1 The quick brown fox jumps over the lazy dog, then naps in the warm afternoon sun.
English M2 Could you remind me to call my sister tomorrow morning? I keep forgetting.
Spanish F2 Hoy hace un día precioso. ¿Te apetece dar un paseo por el parque?
Spanish M1 La biblioteca municipal abre a las nueve y cierra a las ocho de la tarde.
Portuguese F3 Que bom te ver de novo! Vamos tomar um café e conversar um pouco?
Portuguese M3 O comboio para Lisboa parte daqui a vinte minutos, não se atrase.
French F4 Bonjour à tous, et bienvenue dans cette nouvelle émission consacrée à la science.
French M4 Je pense qu'il va pleuvoir ce soir, n'oublie pas ton parapluie.
German F5 Guten Morgen! Hast du gut geschlafen? Heute wird ein langer, aber schöner Tag.
German M5 Die Bahn hat leider zwanzig Minuten Verspätung, wir sollten ein Taxi nehmen.
Italian F1 Che bella giornata! Andiamo a prendere un gelato in piazza?
Italian M2 Il museo resterà chiuso per lavori fino alla fine del mese prossimo.
Swedish F2 Hej! Vill du följa med och fika på det nya kaféet vid torget?
Swedish M1 Tåget mot Göteborg är tyvärr försenat på grund av ett signalfel.
Vietnamese F3 Xin chào các bạn, hôm nay trời thật đẹp, chúng ta cùng đi dạo công viên nhé!
Vietnamese M3 Cuốn sách này kể về hành trình của một cậu bé đi tìm ước mơ của mình.
Japanese F4 こんにちは。今日はとても良い天気ですね。一緒に散歩に行きませんか?
Japanese M4 駅までの道を教えていただけますか?初めてこの町に来ました。
Korean F5 안녕하세요! 오늘 날씨가 정말 좋네요. 같이 산책하러 갈까요?
Korean M5 이번 주말에는 가족들과 함께 바닷가에 놀러 갈 예정이에요.
Indonesian F1 Selamat pagi semuanya! Hari ini cuacanya cerah sekali, ayo kita jalan-jalan.
Indonesian M2 Kereta menuju Bandung akan berangkat sepuluh menit lagi dari peron tiga.

Setup

git clone https://huggingface.co/polyskill/LGTM
cd LGTM
pip install -r requirements.txt          # PyTorch backend
pip install -r requirements-onnx.txt     # ONNX backend

PyTorch

from lgtm import LGTMTTS

tts = LGTMTTS.from_pretrained(".")            # or "polyskill/LGTM" to download
wav = tts.synthesize("Xin chào, hôm nay trời đẹp quá!", lang="vi", voice="F1")
tts.save_wav(wav, "out.wav")                   # 44.1 kHz mono

ONNX Runtime

from lgtm import LGTMOnnx

tts = LGTMOnnx.from_pretrained(".", use_gpu=False)   # use_gpu=True with onnxruntime-gpu
wav = tts.synthesize("Bonjour à tous, comment allez-vous ?", lang="fr", voice="M1")
tts.save_wav(wav, "out.wav")

Voice cloning

Give 5-15 seconds of clean speech; the voice can then speak any supported language.

voice = tts.clone_voice("reference.wav")      # works with both backends
wav = tts.synthesize("This is my cloned voice.", lang="en", voice=voice)

from lgtm import save_voice_style              # ONNX: from lgtm.onnx_inference import save_voice_style
save_voice_style("my_voice.json", voice)       # reuse later: voice="my_voice.json"

Command line

python -m lgtm.cli --text "Hej! Hur mår du idag?" --lang sv --voice F2 --out out.wav
python -m lgtm.cli --text "안녕하세요" --lang ko --ref reference.wav --save_voice my_voice.json --out out.wav
python -m lgtm.cli --text "Selamat pagi" --lang id --voice M3 --backend onnx --out out.wav

Options

argument default
lang "en" language code (see above)
voice "F1" preset name, path to a voice .json, or clone_voice() output
steps 8 denoising steps (fewer = faster, e.g. 5)
speed 1.05 speaking rate (higher = faster)
silence 0.3 seconds of silence between sentences (long text is split automatically)

Files

path contents
pytorch/model.safetensors all weights (synthesis + voice cloning)
onnx/text_encoder.onnx, duration_predictor.onnx, vector_estimator.onnx, vocoder.onnx synthesis graphs
onnx/voice_encoder.onnx reference audio → voice style (cloning)
voice_styles/*.json built-in voices
config.json, unicode_indexer.json model config, text vocabulary
lgtm/ inference code (inference.py PyTorch, onnx_inference.py ONNX, cli.py)

ONNX graph I/O (for custom runtimes)

graph inputs outputs
text_encoder text_ids int64 [B,T], style_ttl [B,50,256], text_mask [B,1,T] text_emb [B,256,T]
duration_predictor text_ids, style_dp [B,8,16], text_mask duration [B] (seconds)
vector_estimator noisy_latent [B,144,L], text_emb, style_ttl, latent_mask [B,1,L], text_mask, current_step [B], total_step [B] denoised_latent [B,144,L]
vocoder latent [B,144,L] wav [B, 3072·L]
voice_encoder wav [1,N] (44.1 kHz) style_ttl [1,50,256], style_dp [1,8,16]

Sampling loop: L = ceil(duration·44100 / 3072), start from Gaussian noise masked by latent_mask, call vector_estimator for current_step = 0 … total_step-1, then vocoder. See lgtm/onnx_inference.py.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support