AudacieuseAi's picture
Initial VibeVoice-Hindi test suite
9be564c verified
|
Raw
History Blame Contribute Delete
3.89 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: VibeVoice Hindi Test
emoji: 🎤
colorFrom: yellow
colorTo: red
sdk: gradio
sdk_version: 5.50.0
app_file: app.py
hardware: zero-a10g
pinned: false

VibeVoice-Hindi-1.5B Test Suite

Test and evaluate the fine-tuned Hindi text-to-speech model from HuggingFace.

Quick Start

Option 1: Deploy to HF Spaces (Recommended)

  1. Create a new HF Space:

  2. Clone this repo into the Space:

    git clone https://huggingface.co/spaces/YOUR-USERNAME/vibevoice-hindi-test
    cd vibevoice-hindi-test
    
  3. Push files:

    git add .
    git commit -m "Initial VibeVoice-Hindi test suite"
    git push
    
  4. HF Spaces will auto-build and deploy (takes ~5 min)

    • Access at: https://huggingface.co/spaces/YOUR-USERNAME/vibevoice-hindi-test

Option 2: Run Locally

pip install -r requirements.txt
python app.py

Then open http://localhost:7860 in your browser.


Test Cases Included

Test Input Purpose
Greeting नमस्ते, आपका स्वागत है। आज आपका दिन कैसा चल रहा है? Test casual speech, daily conversation
Business हमारी कंपनी गुणवत्ता और नवाचार में विश्वास करती है। Test formal tone, business terminology
Numbers साल २०२६ में हमने बीस लाख रुपये की आय प्राप्त की। Test numeral pronunciation, decimal points
Mixed आप YouTube पर हमारा नया वीडियो देख सकते हैं। Test English words mixed with Hindi

Evaluation Rubric

After generating speech, rate on 5 dimensions:

  • 📖 Understand — Can you follow what's being said?
  • 🎵 Natural — Does it sound like a real person (not robot)?
  • ⚡ Speed — Is the inference latency acceptable (<5s)?
  • 🔊 Clarity — Can you hear each word clearly?
  • 🎯 Consistency — Same pronunciation throughout utterance?

Scoring

  • 5/5 ✓ → 🟢 Production Ready
  • 4/5 ✓ → 🟡 Good, Minor Issues
  • 3/5 ✓ → 🟠 Acceptable, Needs Work
  • 2/5 ✓ → 🔴 Poor Quality
  • 1/5 ✓ → ❌ Archive This Model

Evaluation Results

All evaluations are saved to evaluation_log.jsonl with:

  • Timestamp
  • Test case name
  • Dimensions passed (0-5)
  • Optional notes
  • Device used (CPU/GPU)

Model Info

  • Model ID: tarun7r/vibevoice-hindi-1.5B
  • Base: Qwen2.5-1.5B (text understanding)
  • Voice: Diffusion-based speech generation (~600M params)
  • Fine-tuning: LoRA on LLM + full diffusion head tuning
  • License: MIT
  • Size: 3.5GB (fits T4 GPU, ZeroGPU Spaces)

Troubleshooting

Q: "CUDA out of memory"

  • Reduce batch size or use CPU inference (slower but works on Spaces)

Q: Audio generation is placeholder only

Q: Model not loading

  • Check internet connection (HF model download)
  • Verify HF token if model is private (use huggingface-cli login)

Next Steps

  1. ✅ Run the 4 test cases
  2. ✅ Listen and evaluate each dimension
  3. ✅ Take notes on pronunciation issues
  4. ✅ Save evaluations
  5. 📊 Review evaluation_log.jsonl for patterns
  6. 🎯 Decide: Production (1.5B) or test 7B variant?

Estimated time: 15-20 minutes for full evaluation suite