[Feature Request] PersonaPlex-7B: Fine-Tuning, RAG & Custom Voice Support

#23
by mahin2110 - opened

Hi πŸ‘‹ I’m using nvidia/personaplex-7b-v1 to build a real-time voice customer support agent and have a few questions:

  1. Fine-tuning

Are fine-tuning scripts/docs planned?

Supported data formats (audio + transcripts)?

Hardware requirements?

LoRA / QLoRA support?

  1. Custom knowledge integration

Max context/token limit for prompts?

Can context be updated dynamically during a session?

Recommended RAG setup (ASR β†’ RAG β†’ PersonaPlex)?

Latency considerations for real-time/full-duplex use?

Is domain fine-tuning supported/planned?

  1. Custom voice

Custom voice embeddings or voice cloning?

Required audio format and duration?

Current workaround: static prompt-based knowledge
Limitations: token limits, no retrieval, manual updates

Env: A10G 24GB | production voice agent

Thanks! Any guidance or roadmap info would be appreciated.

Hey I have been trying to use the model for some use-cases and want to know is there any official doc or soemthing regarding the fine-tuning of the model.

This comment has been hidden
NVIDIA org

I don't recommend using this model checkpoint for production. Its more of a showcase for naturalness. Stay tuned for future models that are smarter and are packaged with finetuning flows and toolcalling support. Can't promise custom voice though since its extremely difficult to get legal approval for that.

Any updates on this topic?

Sign up or log in to comment