File size: 6,887 Bytes
dd7b1b6
 
 
 
 
 
 
 
 
 
 
194be9c
dd7b1b6
194be9c
 
dd7b1b6
 
194be9c
dd7b1b6
194be9c
 
dd7b1b6
194be9c
dd7b1b6
 
 
 
 
194be9c
 
 
 
 
 
 
 
 
 
 
 
 
66e9754
 
 
 
194be9c
 
 
 
 
 
 
 
 
 
 
 
 
 
dd7b1b6
 
 
 
 
 
 
 
 
 
 
194be9c
 
dd7b1b6
 
 
194be9c
dd7b1b6
 
 
 
194be9c
 
 
dd7b1b6
 
 
 
 
 
 
 
 
 
 
 
 
 
8f81618
dd7b1b6
 
 
8f81618
 
dd7b1b6
 
8f81618
dd7b1b6
 
 
 
 
194be9c
 
 
 
 
 
dd7b1b6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
194be9c
 
 
66e9754
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dd7b1b6
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
---
license: cc-by-4.0
language:
- en
base_model: kyutai/moshiko-pytorch-bf16
datasets:
- ai4bharat/Svarah
tags:
- speech-to-speech
- full-duplex
- spoken-dialogue
- conversational-ai
- voice-agent
- voice-assistant
- real-time
- indian-english
- indian-accent
- india
- customer-support
- call-center
- barge-in
- moshi
- mimi
- lora
- audio
pipeline_tag: audio-to-audio
---

# Roxi-Duplex: a voice agent you can interrupt, in Indian English

Most voice agents take turns: you speak, you wait, they speak. Roxi-Duplex does not wait.
It is a **full-duplex speech-to-speech** model that listens while it talks, so callers can
interrupt it mid-sentence, murmur "haan, okay" while it speaks, and get a response in a
natural **Indian-English** voice that opens with "Welcome to Voz Vox. My name is Roxi.
How may I help you today?"

To our knowledge this is one of the first openly released Indian-English adaptations of a
full-duplex speech-to-speech model. It is a compact LoRA adapter (370 MB) for Kyutai's
Moshi 7B, so you get frontier duplex behavior plus an Indian support-agent persona without
downloading a new foundation model.

Paper: "Roxi-Duplex: Low-Resource Indian-English Adaptation of a Full-Duplex
Speech-to-Speech Model with Synthetic Two-Channel Training Data"
(https://doi.org/10.5281/zenodo.21445239). See the Citation section below.

## Why you might want this

- **True barge-in.** Duplex is architectural, not a VAD hack: the model tracks both audio
  streams on one timeline, so interruptions and backchannels work the way they do between
  humans.
- **Indian-English out of the box.** Accent and delivery learned from Indian speakers, for
  the hundreds of millions of users that US-accented agents serve poorly.
- **Support-agent behavior built in.** Greetings, bookings, order status, complaints,
  payments. Prompt it with a caller and it answers like a call-center agent, not a chatbot.
- **Real time on one GPU.** Moshi streams at a real-time factor of about 0.34 on an A100
  40 GB, with about 200 ms theoretical latency. The adapter adds nothing to inference cost.
- **A reproducible recipe.** Open Indian full-duplex dialogue data does not exist, so we
  synthesized it. The full pipeline is documented below and cheap to rerun (training takes
  17 minutes on one A100).

## Quick facts

| Field | Value |
|---|---|
| Base model | kyutai/moshiko-pytorch-bf16 (7B, CC-BY 4.0) |
| Audio codec | Mimi (streaming, 12.5 Hz frames, 24 kHz) |
| Method | LoRA rank 64, scaling 2.0, 1500 steps, via Kyutai moshi-finetune |
| Adapter size | 370 MB (safetensors) |
| Training data | 150 synthetic two-channel support conversations, about 94 minutes |
| Assistant channel | Roxi TTS (Indian-English, 1.7B MOSS-TTS-Local fine-tune) speaking scripted support dialogues |
| User channel | Real Indian-English speakers from ai4bharat/Svarah (117 speakers, CC-BY 4.0) |
| Training cost | About 17 minutes on one rented A100 40 GB |

## How the data was made

There is no open Indian-English full-duplex dialogue corpus, so we built one:

1. Scripted VozVox support dialogues (greetings, bookings, complaints, payments) with Indian
   names, cities, and numbers written as words.
2. Assistant turns rendered with an Indian-English TTS voice (Roxi), silence-trimmed and
   time-stretched 1.25x with WSOLA for a natural pace. WSOLA matters: phase-vocoder
   stretching made the voice sound robotic.
3. User turns taken from real Svarah recordings across India.
4. Both sides placed on a shared stereo timeline with turn gaps, backchannels, and overlaps,
   plus word-level alignments for Moshi's inner-monologue text stream.

## Usage

Requires a Linux GPU with the `moshi` package (Triton is needed for the real-time compiled
path, so native Windows is not supported for real-time use).

```bash
pip install moshi
```

```python
import torch
from huggingface_hub import hf_hub_download
from moshi.models.loaders import CheckpointInfo, get_lora_moshi
from moshi.models import LMGen

adapter = hf_hub_download("IOTEverythin/roxi-duplex", "lora.safetensors")
info = CheckpointInfo.from_hf_repo("kyutai/moshiko-pytorch-bf16")
mimi = info.get_mimi(device="cuda")
lm = info.get_moshi(device="cuda", dtype=torch.bfloat16)
lm = get_lora_moshi(lm, adapter, 64, 2.0,
                    dtype=torch.bfloat16, device="cuda", fuse_lora=True)
lm_gen = LMGen(lm, temp=0.7, temp_text=0.7)
# Stream user audio through mimi.encode and lm_gen.step exactly as with base Moshi.
```

Two tips from our experiments:

- Prompt the model with real user audio. Feeding only silence makes any Moshi-family model
  produce unfocused speech.
- If you fine-tune further, stay light. Around rank 64 and 1500 steps was the sweet spot;
  heavier adapters (rank 96, 3000 steps) kept the accent but degraded intelligibility.

## Limitations

- Proof-of-concept scale: 150 synthetic conversations from four scripted scenario templates.
  Coverage outside customer-support topics is limited.
- The dialogue structure is synthetic; real-call turn-taking dynamics may differ.
- English only (Indian-English accent); no Hindi code-switching yet.
- Inherits all Moshi limitations and requires a GPU for real-time use.

## License and attribution

Released under **CC-BY 4.0**, matching the base model. This work builds on:

- **Moshi and Mimi** by Kyutai (kyutai/moshiko-pytorch-bf16, CC-BY 4.0). Defossez et al.,
  "Moshi: a speech-text foundation model for real-time dialogue".
- **Svarah** by AI4Bharat (ai4bharat/Svarah, CC-BY 4.0), used for the user audio channel.
- The assistant voice derives from our Roxi TTS, fine-tuned on the IIT-Madras Indic TTS English
  set. Required notice: COPYRIGHT 2016 TTS Consortium, TDIL, Meity, represented by Hema A.
  Murthy and S. Umesh, Department of Computer Science and Engineering and Electrical
  Engineering, IIT Madras. ALL RIGHTS RESERVED.
- Trained with Kyutai's moshi-finetune.

See also our Indian-English TTS models: IOTEverythin/roxi-tts-pro (1.7B premium) and
IOTEverythin/roxi-tts-v3.1 (0.1B real-time).

## Citation

If you use this model, the data recipe, or the RoxiDuplex-Eval benchmark, please cite:

```bibtex
@misc{roxiduplex2026,
  title     = {Roxi-Duplex: Low-Resource Indian-English Adaptation of a Full-Duplex
               Speech-to-Speech Model with Synthetic Two-Channel Training Data},
  author    = {A, Joshua Nishanth Tarun and A, Joel Ajitesh Varun},
  year      = {2026},
  month     = {july},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21445239},
  url       = {https://doi.org/10.5281/zenodo.21445239},
  note      = {Preprint}
}
```

## Responsible use

This model speaks with a synthetic voice derived from consented and licensed datasets. Do not
use it to impersonate real people or for fraud, social engineering, or deception. Disclose
AI-generated audio where required by law or policy. Provided as is, without warranty.