Emotion-enhanced TTS via LLMβTTS FFN blending. No training required. Now with Voice Cloning!
LLM FFN blended β more expressive, emotional speech
Upload 3~15 seconds of clear speech audio. The model will clone this voice.
If no audio is uploaded, a default voice will be used.