Indic Smart Turn

A voice agent has to decide, every time the user pauses, whether they have finished speaking or are only taking a breath. Get it wrong one way and the agent interrupts; get it wrong the other way and it sits in silence. Indic Smart Turn makes that decision from the audio itself, for Indic-language conversations (and English), and runs on a single CPU core in under 50 ms.

It is built with the Pipecat Smart Turn v3 recipe, not from its weights. The encoder starts from OpenAI's Whisper (base or tiny), the decoder is dropped, Smart Turn's attention-pooling head is added, and the whole network is trained on 8-second complete-or-incomplete clips: about 52,000 from real Indian phone conversations, plus Pipecat's own data so English stays in the mix. Smart Turn v3.2 appears here only as the model we compare against. The result loads in Pipecat's LocalSmartTurnAnalyzerV3 with a one-line change.

Why this model

  • Covers the languages Smart Turn v3.2 does not. Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Odia and Assamese are absent from v3.2. Hindi, Marathi and Bengali are present there but were trained mostly on synthetic speech.
  • Trained on real speech. The Indian-language training clips are real two-person phone conversations recorded by AI4Bharat for IndicVoices, not synthetic TTS. Labels come from an audio model listening to each clip, with a text model as a second opinion.
  • Measured against the same test clips as Smart Turn v3.2. Ahead by 3 to 14 points at the same model size, and by 6 to 18 points with the larger encoder. On the human-validated TamilEOT benchmark it matches the published whisper-base result.
  • One file for all languages. No per-language switching, and English is kept.

Languages

  • English
  • Hindi
  • Marathi
  • Bengali
  • Tamil
  • Telugu
  • Kannada
  • Malayalam
  • Gujarati
  • Punjabi
  • Odia
  • Assamese

Which file to use

Recommended: indic-smart-turn-base-int8.onnx on x86 servers. Use indic-smart-turn-base-fp32.onnx on Apple Silicon or GPU hosts, and indic-smart-turn-tiny-int8.onnx on constrained CPUs.

indic-smart-turn-base-int8.onnx

  • Version 1.1: tuned on production calls (see "Fine-tuning on production data"); the previous file is indic-smart-turn-base-int8-v1.onnx.
  • Encoder: whisper-base. Precision: int8, dynamic (weights only). Size: 24 MB.
  • Latency, single thread, batch 1: 122 ms on a server core, 44 ms on Apple Silicon.
  • ROC-AUC within 0.005 of the fp32 model in every language; accuracy 1 to 2.6 points lower at the 0.5 threshold.
  • Best accuracy per millisecond on x86.

indic-smart-turn-base-fp32.onnx

  • Version 1.1: tuned on production calls; the previous file is indic-smart-turn-base-fp32-v1.onnx.
  • Encoder: whisper-base. Precision: fp32. Size: 81 MB.
  • Latency: 190 to 208 ms on a server core, 37 ms on Apple Silicon.
  • The most accurate model: 84 to 95% accuracy, AUC 0.86 to 0.99. On TamilEOT it scores 85.9%, matching the paper's whisper-base result of 86.1%.
  • On Apple Silicon and GPUs it runs as fast as tiny, so it is the better choice there.

indic-smart-turn-tiny-int8.onnx

  • Encoder: whisper-tiny. Precision: int8, static. Size: 8.8 MB.
  • Latency: 83 ms on a server core, 36 ms on Apple Silicon. Same size and speed as Smart Turn v3.2's CPU file.
  • Ahead of Smart Turn v3.2 int8 in every Indian language by 3 to 14 points; 1.2 points behind on English.

indic-smart-turn-tiny-fp32.onnx

  • Encoder: whisper-tiny. Precision: fp32. Size: 32 MB.
  • Latency: about the same as base int8 on a server core, 36 ms on Apple Silicon.
  • 1 to 4 points above tiny int8. Worth it only where fp32 costs nothing extra.

Results

Accuracy on a speaker-disjoint test split of 20,431 clips at the default 0.5 threshold; each chart row also shows ROC-AUC, which does not depend on the threshold. Smart Turn v3.2 is Pipecat's shipped model, a whisper-tiny in two files: int8 for CPU and fp32 for GPU. There is no base-size v3.2, so each group is compared with the v3.2 file of the same precision.

  • Test data: IndicVoices test clips for the eleven Indian languages; English, Hindi, Marathi and Bengali also include Pipecat's own v3.2 test set; Tamil also includes the human-validated TamilEOT test set.
  • Full tables with ROC-AUC and 95% bootstrap intervals: eval_all_test.md. Discussion: STAGE6_REPORT.md. Interactive version of the charts with hover values and table views: charts.html in this repo.

Group 1: tiny int8 vs Smart Turn v3.2 int8

tiny int8 vs Smart Turn v3.2 int8

Group 2: base int8 vs Smart Turn v3.2 int8

base int8 vs Smart Turn v3.2 int8

Group 3: tiny fp32 vs Smart Turn v3.2 fp32

tiny fp32 vs Smart Turn v3.2 fp32

Group 4: base fp32 vs Smart Turn v3.2 fp32

base fp32 vs Smart Turn v3.2 fp32

External test sets

Set Smart Turn v3.2 int8 tiny int8 base int8 base fp32
TamilEOT test, 4,168 human-validated clips 70.4 / 0.743 81.8 / 0.875 84.8 / 0.917 85.9 / 0.917
Pipecat v3.2 test, eng/hin/mar/ben, 10,878 clips 90.6 / 0.968 90.2 / 0.966 93.6 / 0.983

Values are accuracy / ROC-AUC. The TamilEOT paper reports 86.1 for its own whisper-base fine-tune.

Reading the numbers

The model gives each pause a probability that the turn is complete. We call it "complete" when that probability is above 0.5. Accuracy counts how many of those calls were right. ROC-AUC asks a different question: if you hand the model one complete pause and one incomplete pause, how often does it give the complete one the higher probability? AUC ignores where the 0.5 line sits; accuracy depends on it.

Why Smart Turn v3.2 int8 scores higher than Smart Turn v3.2 fp32 on accuracy in 11 of 12 languages. The two files are the same model at two precisions, and on AUC they are within 0.03 of each other everywhere, so they tell the two kinds of pause apart equally well. What differs is where their probabilities sit. Quantisation pushed the int8 file's probabilities up, so it says "complete" more readily; the fp32 file says "incomplete" more readily. In these test sets 57 to 81% of pauses are complete, so a model that says "complete" more often gets more calls right at the 0.5 line. The int8 file is not a better model; its probabilities just happen to sit on the better side of the line for this data. Shift the line, and the ordering can flip.

The table shows this on six languages: the AUC columns are close, while the fp32 file catches far fewer complete pauses at 0.5.

Language int8 accuracy fp32 accuracy int8 AUC fp32 AUC complete pauses caught at 0.5, fp32 same, int8
Kannada 77.7 56.6 0.770 0.795 50.0% 81.9%
Malayalam 66.9 49.6 0.657 0.667 39.1% 74.7%
Telugu 73.4 58.1 0.750 0.773 46.5% 79.6%
Gujarati 73.5 58.4 0.765 0.759 49.9% 78.6%
Tamil 70.5 58.6 0.744 0.744 46.1% 76.0%
English 92.9 94.7 0.978 0.986 95.3% 93.6%

Two consequences for reading the charts:

  • The fp32 groups show bigger accuracy gaps than the int8 groups partly because the v3.2 fp32 file's 0.5 line is badly placed for this data, not only because our fp32 models are stronger. The AUC printed at the end of each row is the fairer comparison.
  • Our own int8 and fp32 files stay close to each other: base int8 and base fp32 land on the same side of the 0.5 line for 95.9% of test clips, with a median probability difference of 0.001, so switching precision does not move the operating point the way it does for Smart Turn v3.2. A few lower-precision or smaller variants still post a slightly higher accuracy than their bigger sibling: tiny int8 over tiny fp32 on Telugu (+0.1), Kannada (+0.6) and Odia (+1.5), and base int8 over base fp32 on Odia (+0.6). These are the same effect at a much smaller scale, and all of them are inside the ±3 point confidence intervals of those test sets. On AUC, base beats tiny in every language, and fp32 is equal to or above int8 for base in every language. Treat accuracy gaps under about 2 points on the smaller languages (331 to 725 test clips) as ties.

Known limits

  • Weakest languages: Assamese, Gujarati and Malayalam, at 82 to 85% with base fp32.
  • All files use the default threshold of 0.5.

Production calls

The models were also tested on 1,200 real roleplay sessions from a voice-agent product: a trainee practising a sales pitch with a TTS agent, mixed into one mono track per session. This data is not published, in any form; only the aggregate numbers below leave the machine.

How the test set was made

  • Each session was split into the trainee's and the agent's speech by voice. Sessions where the two could not be separated with confidence were excluded, about a quarter of them.
  • Every pause of 200 ms or more in the trainee's speech became a test clip: the last 8 s of trainee-only audio ending 0.2 s after the pause, the same cut as the training data.
  • Labels are gemini-3.7-flash audio verdicts, the same labeler as the public test set, so the two evaluations are directly comparable. 12,710 labelled test clips across ten languages, split by session.

Accuracy / ROC-AUC at the 0.5 threshold

Production calls: Indic base int8 vs Smart Turn v3.2 int8

Language Test clips Smart Turn v3.2 int8 Indic tiny int8 Indic base int8
Tamil 991 61.9 / 0.755 70.5 / 0.819 78.7 / 0.875
Malayalam 462 64.1 / 0.732 75.8 / 0.819 78.4 / 0.862
Bengali 2116 64.7 / 0.769 70.3 / 0.814 77.5 / 0.859
Telugu 1097 66.2 / 0.788 71.8 / 0.825 77.3 / 0.886
Marathi 925 66.4 / 0.780 73.4 / 0.820 77.0 / 0.854
Odia * 95 70.5 / 0.810 66.3 / 0.860 78.9 / 0.912
Hindi 2165 68.9 / 0.797 75.0 / 0.837 78.8 / 0.876
Kannada 367 64.3 / 0.778 67.8 / 0.819 74.1 / 0.864
Gujarati 2006 68.6 / 0.770 76.2 / 0.840 76.5 / 0.857
English 2486 67.6 / 0.752 71.2 / 0.794 75.2 / 0.829
All 12710 66.6 / 0.771 72.7 / 0.821 77.1 / 0.859

* Odia: 95 clips from 2 sessions, indicative only.

Smart Turn v3.2 int8 Indic tiny int8 Indic base int8
Recall on complete turns 84.5% 84.5% 85.7%
Recall on incomplete turns 52.8% 63.6% 70.4%

What it shows

  • Indic Smart Turn is ahead in every language, by 8 to 18 points of accuracy and 0.06 to 0.13 of AUC with the base int8 model, and by 4 to 12 points with tiny int8, which is the same size and speed as Smart Turn v3.2.
  • The gain sits where it did on the public data: both systems recognise a finished turn about equally (85% recall), ours is far better at recognising an unfinished one (71% against 53%).
  • All numbers are lower than on the public test set because a sales pitch has many more mid-sentence and sentence-final pauses than a conversation; both systems drop by a similar amount, and the gap between them is the result.

Usage

Pipecat

from pipecat.audio.turn.smart_turn.local_smart_turn_v3 import LocalSmartTurnAnalyzerV3

analyzer = LocalSmartTurnAnalyzerV3(smart_turn_model_path="indic-smart-turn-base-int8.onnx")

Direct ONNX

The model takes an 80 x 800 log-mel of the last 8 seconds of 16 kHz mono audio and returns the probability that the turn is complete. config.json records the same contract.

Single clip:

import numpy as np, onnxruntime as ort, soundfile as sf
from transformers import WhisperFeatureExtractor

MODEL = "indic-smart-turn-base-int8.onnx"
fe = WhisperFeatureExtractor(chunk_length=8)           # 8 s window, 80 mel bins, 800 frames
so = ort.SessionOptions(); so.intra_op_num_threads = 1  # one core is enough at batch 1
session = ort.InferenceSession(MODEL, so, providers=["CPUExecutionProvider"])

def last_8s(audio: np.ndarray, sr: int = 16000) -> np.ndarray:
    n = 8 * sr
    return audio[-n:] if len(audio) > n else np.pad(audio, (n - len(audio), 0))  # keep the end, zero-pad the front

def turn_complete_prob(audio: np.ndarray) -> float:
    feats = fe(last_8s(audio), sampling_rate=16000, return_tensors="np", padding="max_length",
               max_length=8 * 16000, truncation=True, do_normalize=True).input_features.astype(np.float32)
    return float(session.run(None, {"input_features": feats})[0][0, 0])

audio, sr = sf.read("clip.wav", dtype="float32")        # 16 kHz mono; resample first if not
p = turn_complete_prob(audio)
print("complete" if p > 0.5 else "incomplete", round(p, 3))

Batch of clips:

def turn_complete_probs(clips: list[np.ndarray], batch_size: int = 64) -> np.ndarray:
    out = []
    for i in range(0, len(clips), batch_size):
        wavs = [last_8s(c) for c in clips[i:i + batch_size]]
        feats = fe(wavs, sampling_rate=16000, return_tensors="np", padding="max_length",
                   max_length=8 * 16000, truncation=True, do_normalize=True).input_features.astype(np.float32)
        out.append(session.run(None, {"input_features": feats})[0][:, 0])
    return np.concatenate(out)

probs = turn_complete_probs([sf.read(f, dtype="float32")[0] for f in ["a.wav", "b.wav", "c.wav"]])

For batch scoring raise intra_op_num_threads to the number of cores you can spare.

Fine-tuning on production data

After the public training run, the base model was tuned further on the production calls.

  • Starting point: the Stage 5 whisper-base checkpoint, not Whisper weights.
  • Data: 20,649 production clips where the recording and Gemini agree (4,601 complete, 16,048 incomplete), mixed one-to-one with 20,000 clips sampled from the public Indic training set, plus 3,000 TamilEOT and 5,000 English clips. 48,649 clips in total.
  • Recipe: one pass over the data, learning rate 1e-5 (five times smaller than the 5e-5 used for the original training), batch 128, otherwise unchanged. 7.8 minutes on one RTX A6000.
  • Threshold: the tuned model's probabilities sit lower than the original's, so its best decision line is 0.13 rather than 0.5. The line was chosen on the production train split only, then baked into the exported graph as a constant shift before the final sigmoid, so the shipped files work at the standard 0.5 and the ranking of clips is unchanged.
  • Export: fp32 ONNX and dynamic int8, as for the released models.

Result, int8 files, accuracy / ROC-AUC:

Test set Before tuning After tuning
Production calls, 12,710 clips 78.0 / 0.855 80.0 / 0.876
Public test, 20,431 clips 88.6 / 0.953 90.0 / 0.957

The tuned model gained on both and is ahead in 11 of 12 public languages (Odia, 331 clips, is 2.7 points lower). It replaces the original base files on Hugging Face as indic-smart-turn-base-int8.onnx and indic-smart-turn-base-fp32.onnx; the original files remain as indic-smart-turn-base-int8-v1.onnx and indic-smart-turn-base-fp32-v1.onnx. Full tables are in the project repository under reports/.

Training

  • Architecture: Whisper encoder (decoder discarded), attention pooling, small MLP head. The upstream Smart Turn v3 layout, unchanged. Weights start from openai/whisper-base or openai/whisper-tiny; Smart Turn v3.2's weights are never used as a starting point.
  • Data per epoch: 166,939 samples. IndicVoices conversational clips for the eleven languages (51,917, oversampled twice), the TamilEOT train split, and Pipecat's v3.2 data for English (capped at 40,000), Hindi, Marathi and Bengali.
  • Recipe: upstream trainer, from Whisper weights, 4 epochs, batch 128, lr 5e-5, cosine schedule with 20% warmup, bf16, best checkpoint by eval F1.
  • Hardware: one NVIDIA RTX A6000 on RunPod. 19 minutes for base, 24 for tiny.
  • Labels for the IndicVoices clips: audio verdicts from gemini-3.7-flash, the labeler the TamilEOT paper validated at 97.5% agreement with human listeners, with a text-LLM second opinion that agrees 89 to 93% of the time. Test clips where the two disagree are held out as ambiguous.
  • Full method, data card and code: dataset adimyth/indic-smart-turn-data and the project repository.

License and attribution

Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adimyth/indic-smart-turn

Quantized
(249)
this model

Datasets used to train adimyth/indic-smart-turn

Paper for adimyth/indic-smart-turn