🧬 FINAL-Bench · World's First Cross-Modal FFN Transfer

Darwin-TTS-1.7B-Cross v2

Emotion-enhanced TTS via LLM→TTS FFN blending. No training required. Now with Voice Cloning!

🧬
84
FFN Tensors
βœ…
100%
Shape Match
⚑
0
Training
πŸŽ™οΈ
Clone
Voice Cloning

🧬 Darwin Enhancement

LLM FFN blended β€” more expressive, emotional speech

πŸŽ™οΈ Voice Cloning (Optional)

Upload 3~15 seconds of clear speech audio. The model will clone this voice.
If no audio is uploaded, a default voice will be used.

🎀
Click to upload or drag & drop
WAV, MP3, FLAC Β· 3~15 seconds recommended
🧬 Generating speech...
This may take 30~60 seconds on first run.

✨ Darwin-TTS Output

πŸ”¬ How It Works

Darwin-TTS v2 Architecture:
β”œβ”€β”€ talker (28L Qwen3 LM) ← Layer-selective FFN blending (proprietary)
β”œβ”€β”€ code_predictor (5L) ← untouched
β”œβ”€β”€ speech_tokenizer ← untouched
└── encoder/decoder ← untouched

v2: Long text stability + quality exceeds original + Voice Cloning