roxi-duplex / README.md
Joshuant's picture
add paper citation (Zenodo DOI 10.5281/zenodo.21445239)
66e9754 verified
|
Raw
History Blame Contribute Delete
6.89 kB
---
license: cc-by-4.0
language:
- en
base_model: kyutai/moshiko-pytorch-bf16
datasets:
- ai4bharat/Svarah
tags:
- speech-to-speech
- full-duplex
- spoken-dialogue
- conversational-ai
- voice-agent
- voice-assistant
- real-time
- indian-english
- indian-accent
- india
- customer-support
- call-center
- barge-in
- moshi
- mimi
- lora
- audio
pipeline_tag: audio-to-audio
---
# Roxi-Duplex: a voice agent you can interrupt, in Indian English
Most voice agents take turns: you speak, you wait, they speak. Roxi-Duplex does not wait.
It is a **full-duplex speech-to-speech** model that listens while it talks, so callers can
interrupt it mid-sentence, murmur "haan, okay" while it speaks, and get a response in a
natural **Indian-English** voice that opens with "Welcome to Voz Vox. My name is Roxi.
How may I help you today?"
To our knowledge this is one of the first openly released Indian-English adaptations of a
full-duplex speech-to-speech model. It is a compact LoRA adapter (370 MB) for Kyutai's
Moshi 7B, so you get frontier duplex behavior plus an Indian support-agent persona without
downloading a new foundation model.
Paper: "Roxi-Duplex: Low-Resource Indian-English Adaptation of a Full-Duplex
Speech-to-Speech Model with Synthetic Two-Channel Training Data"
(https://doi.org/10.5281/zenodo.21445239). See the Citation section below.
## Why you might want this
- **True barge-in.** Duplex is architectural, not a VAD hack: the model tracks both audio
streams on one timeline, so interruptions and backchannels work the way they do between
humans.
- **Indian-English out of the box.** Accent and delivery learned from Indian speakers, for
the hundreds of millions of users that US-accented agents serve poorly.
- **Support-agent behavior built in.** Greetings, bookings, order status, complaints,
payments. Prompt it with a caller and it answers like a call-center agent, not a chatbot.
- **Real time on one GPU.** Moshi streams at a real-time factor of about 0.34 on an A100
40 GB, with about 200 ms theoretical latency. The adapter adds nothing to inference cost.
- **A reproducible recipe.** Open Indian full-duplex dialogue data does not exist, so we
synthesized it. The full pipeline is documented below and cheap to rerun (training takes
17 minutes on one A100).
## Quick facts
| Field | Value |
|---|---|
| Base model | kyutai/moshiko-pytorch-bf16 (7B, CC-BY 4.0) |
| Audio codec | Mimi (streaming, 12.5 Hz frames, 24 kHz) |
| Method | LoRA rank 64, scaling 2.0, 1500 steps, via Kyutai moshi-finetune |
| Adapter size | 370 MB (safetensors) |
| Training data | 150 synthetic two-channel support conversations, about 94 minutes |
| Assistant channel | Roxi TTS (Indian-English, 1.7B MOSS-TTS-Local fine-tune) speaking scripted support dialogues |
| User channel | Real Indian-English speakers from ai4bharat/Svarah (117 speakers, CC-BY 4.0) |
| Training cost | About 17 minutes on one rented A100 40 GB |
## How the data was made
There is no open Indian-English full-duplex dialogue corpus, so we built one:
1. Scripted VozVox support dialogues (greetings, bookings, complaints, payments) with Indian
names, cities, and numbers written as words.
2. Assistant turns rendered with an Indian-English TTS voice (Roxi), silence-trimmed and
time-stretched 1.25x with WSOLA for a natural pace. WSOLA matters: phase-vocoder
stretching made the voice sound robotic.
3. User turns taken from real Svarah recordings across India.
4. Both sides placed on a shared stereo timeline with turn gaps, backchannels, and overlaps,
plus word-level alignments for Moshi's inner-monologue text stream.
## Usage
Requires a Linux GPU with the `moshi` package (Triton is needed for the real-time compiled
path, so native Windows is not supported for real-time use).
```bash
pip install moshi
```
```python
import torch
from huggingface_hub import hf_hub_download
from moshi.models.loaders import CheckpointInfo, get_lora_moshi
from moshi.models import LMGen
adapter = hf_hub_download("IOTEverythin/roxi-duplex", "lora.safetensors")
info = CheckpointInfo.from_hf_repo("kyutai/moshiko-pytorch-bf16")
mimi = info.get_mimi(device="cuda")
lm = info.get_moshi(device="cuda", dtype=torch.bfloat16)
lm = get_lora_moshi(lm, adapter, 64, 2.0,
dtype=torch.bfloat16, device="cuda", fuse_lora=True)
lm_gen = LMGen(lm, temp=0.7, temp_text=0.7)
# Stream user audio through mimi.encode and lm_gen.step exactly as with base Moshi.
```
Two tips from our experiments:
- Prompt the model with real user audio. Feeding only silence makes any Moshi-family model
produce unfocused speech.
- If you fine-tune further, stay light. Around rank 64 and 1500 steps was the sweet spot;
heavier adapters (rank 96, 3000 steps) kept the accent but degraded intelligibility.
## Limitations
- Proof-of-concept scale: 150 synthetic conversations from four scripted scenario templates.
Coverage outside customer-support topics is limited.
- The dialogue structure is synthetic; real-call turn-taking dynamics may differ.
- English only (Indian-English accent); no Hindi code-switching yet.
- Inherits all Moshi limitations and requires a GPU for real-time use.
## License and attribution
Released under **CC-BY 4.0**, matching the base model. This work builds on:
- **Moshi and Mimi** by Kyutai (kyutai/moshiko-pytorch-bf16, CC-BY 4.0). Defossez et al.,
"Moshi: a speech-text foundation model for real-time dialogue".
- **Svarah** by AI4Bharat (ai4bharat/Svarah, CC-BY 4.0), used for the user audio channel.
- The assistant voice derives from our Roxi TTS, fine-tuned on the IIT-Madras Indic TTS English
set. Required notice: COPYRIGHT 2016 TTS Consortium, TDIL, Meity, represented by Hema A.
Murthy and S. Umesh, Department of Computer Science and Engineering and Electrical
Engineering, IIT Madras. ALL RIGHTS RESERVED.
- Trained with Kyutai's moshi-finetune.
See also our Indian-English TTS models: IOTEverythin/roxi-tts-pro (1.7B premium) and
IOTEverythin/roxi-tts-v3.1 (0.1B real-time).
## Citation
If you use this model, the data recipe, or the RoxiDuplex-Eval benchmark, please cite:
```bibtex
@misc{roxiduplex2026,
title = {Roxi-Duplex: Low-Resource Indian-English Adaptation of a Full-Duplex
Speech-to-Speech Model with Synthetic Two-Channel Training Data},
author = {A, Joshua Nishanth Tarun and A, Joel Ajitesh Varun},
year = {2026},
month = {july},
publisher = {Zenodo},
doi = {10.5281/zenodo.21445239},
url = {https://doi.org/10.5281/zenodo.21445239},
note = {Preprint}
}
```
## Responsible use
This model speaks with a synthetic voice derived from consented and licensed datasets. Do not
use it to impersonate real people or for fraud, social engineering, or deception. Disclose
AI-generated audio where required by law or policy. Provided as is, without warranty.