Thiva-indic by Cortiqa (FP16 Studio Quality)

Thiva-indic is a high-fidelity, on-device Text-to-Speech (TTS) model optimized for the Cortiqa Voice Studio ecosystem. It provides clear, natural, and expressive synthetic speech across 17 Indian languages and English.

The model is optimized in FP16 half-precision, reducing memory footprint by over 50 percent while preserving original studio acoustics.

Key Features

  • 17 Languages in One Model: Generates speech in 16 major Indian languages plus English without separate voice packs.
  • Studio-Grade Audio: 44.1 kHz clean sampling rate with natural breathing and human-like intonation.
  • On-Device and Offline Ready: Engineered for local edge/mobile devices (Android, iOS, laptops) with zero server dependencies.
  • Free for Commercial Use: Released under the permissive Apache 2.0 License.

Supported Languages

# Language Native Script Code
1 Hindi हिन्दी hi
2 Bengali বাংলা bn
3 Tamil தமிழ் ta
4 Telugu తెలుగు te
5 Marathi मराठी mr
6 Gujarati ગુજરાતી gu
7 Kannada ಕನ್ನಡ kn
8 Malayalam മലയാളം ml
9 Odia ଓଡ଼ିଆ or
10 Punjabi ਪੰਜਾਬੀ pa
11 Assamese অসমীয়া as
12 Urdu اردو ur
13 Kashmiri कॉशुर / کٲشُر ks
14 Nepali नेपाली ne
15 Sanskrit संस्कृतम् sa
16 Sindhi سنڌي sd
17 English English en

Quickstart and Inference

import torch
import soundfile as sf
from parler_tts import ParlerTTSForConditionalGeneration
from transformers import AutoTokenizer

device = 'cuda:0' if torch.cuda.is_available() else 'cpu'
model = ParlerTTSForConditionalGeneration.from_pretrained('cortiqa/Thiva-indic').to(device).half()
tokenizer = AutoTokenizer.from_pretrained('cortiqa/Thiva-indic')
description_tokenizer = AutoTokenizer.from_pretrained(model.config.text_encoder._name_or_path)

prompt = 'नमस्ते! कोर्टिका के थिवा-इंडिक मॉडल में आपका स्वागत है।'
description = 'Divya speaks with a clear, pleasant, natural and articulate voice at a steady pace in a quiet studio.'

desc_inputs = description_tokenizer(description, return_tensors='pt').to(device)
prompt_inputs = tokenizer(prompt, return_tensors='pt').to(device)

with torch.no_grad():
    generation = model.generate(
        input_ids=desc_inputs.input_ids,
        attention_mask=desc_inputs.attention_mask,
        prompt_input_ids=prompt_inputs.input_ids,
        prompt_attention_mask=prompt_inputs.attention_mask,
        temperature=0.8,
        do_sample=True,
    )

audio_arr = generation.cpu().float().numpy().squeeze()
sf.write('output.wav', audio_arr, model.config.sampling_rate)
print('Audio generated: output.wav')

Voice Directing and Style Prompts

Control speaker style, gender, pacing, and mood using the description string:

  • Clean Studio Female: 'Divya speaks with a clear, warm, pleasant and articulate voice at a steady pace in a quiet studio with very clean audio.'
  • Punchy Studio Male: 'Rohit speaks with an energetic, modern, confident and clear male voice in a professional recording booth.'
  • Formal News Narration: 'A female speaker delivers an authoritative, clear and precise news narration with deliberate pacing and zero background noise.'

Attribution and Credits

Thiva-indic is built upon and optimized from ai4bharat/indic-parler-tts developed by AI4Bharat (IIT Madras).

License

This model is licensed under the Apache License, Version 2.0. Free to use, modify, distribute, and integrate into commercial applications.

Downloads last month
24
Safetensors
Model size
0.9B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support