Spaces:
Build error
Build error
A newer version of the Gradio SDK is available: 6.22.0
metadata
title: VibeVoice Hindi Test
emoji: 🎤
colorFrom: yellow
colorTo: red
sdk: gradio
sdk_version: 5.50.0
app_file: app.py
hardware: zero-a10g
pinned: false
VibeVoice-Hindi-1.5B Test Suite
Test and evaluate the fine-tuned Hindi text-to-speech model from HuggingFace.
Quick Start
Option 1: Deploy to HF Spaces (Recommended)
Create a new HF Space:
- Go to https://huggingface.co/new-space
- Space name:
vibevoice-hindi-test - License: MIT
- Space SDK: Docker
- Private: No (or Yes if you prefer)
Clone this repo into the Space:
git clone https://huggingface.co/spaces/YOUR-USERNAME/vibevoice-hindi-test cd vibevoice-hindi-testPush files:
git add . git commit -m "Initial VibeVoice-Hindi test suite" git pushHF Spaces will auto-build and deploy (takes ~5 min)
- Access at:
https://huggingface.co/spaces/YOUR-USERNAME/vibevoice-hindi-test
- Access at:
Option 2: Run Locally
pip install -r requirements.txt
python app.py
Then open http://localhost:7860 in your browser.
Test Cases Included
| Test | Input | Purpose |
|---|---|---|
| Greeting | नमस्ते, आपका स्वागत है। आज आपका दिन कैसा चल रहा है? | Test casual speech, daily conversation |
| Business | हमारी कंपनी गुणवत्ता और नवाचार में विश्वास करती है। | Test formal tone, business terminology |
| Numbers | साल २०२६ में हमने बीस लाख रुपये की आय प्राप्त की। | Test numeral pronunciation, decimal points |
| Mixed | आप YouTube पर हमारा नया वीडियो देख सकते हैं। | Test English words mixed with Hindi |
Evaluation Rubric
After generating speech, rate on 5 dimensions:
- 📖 Understand — Can you follow what's being said?
- 🎵 Natural — Does it sound like a real person (not robot)?
- ⚡ Speed — Is the inference latency acceptable (<5s)?
- 🔊 Clarity — Can you hear each word clearly?
- 🎯 Consistency — Same pronunciation throughout utterance?
Scoring
- 5/5 ✓ → 🟢 Production Ready
- 4/5 ✓ → 🟡 Good, Minor Issues
- 3/5 ✓ → 🟠 Acceptable, Needs Work
- 2/5 ✓ → 🔴 Poor Quality
- 1/5 ✓ → ❌ Archive This Model
Evaluation Results
All evaluations are saved to evaluation_log.jsonl with:
- Timestamp
- Test case name
- Dimensions passed (0-5)
- Optional notes
- Device used (CPU/GPU)
Model Info
- Model ID:
tarun7r/vibevoice-hindi-1.5B - Base: Qwen2.5-1.5B (text understanding)
- Voice: Diffusion-based speech generation (~600M params)
- Fine-tuning: LoRA on LLM + full diffusion head tuning
- License: MIT
- Size: 3.5GB (fits T4 GPU, ZeroGPU Spaces)
Troubleshooting
Q: "CUDA out of memory"
- Reduce batch size or use CPU inference (slower but works on Spaces)
Q: Audio generation is placeholder only
- This is a demo scaffold. Integrate actual VibeVoice inference from: https://github.com/tarun7r/vibevoice-tts
Q: Model not loading
- Check internet connection (HF model download)
- Verify HF token if model is private (use
huggingface-cli login)
Next Steps
- ✅ Run the 4 test cases
- ✅ Listen and evaluate each dimension
- ✅ Take notes on pronunciation issues
- ✅ Save evaluations
- 📊 Review
evaluation_log.jsonlfor patterns - 🎯 Decide: Production (1.5B) or test 7B variant?
Estimated time: 15-20 minutes for full evaluation suite