TwiSpeechModel / FIXES_APPLIED.md
ProvLeon's picture
Upload folder using huggingface_hub
6db3515 verified
|
Raw
History Blame Contribute Delete
4.89 kB

A newer version of the Gradio SDK is available: 6.26.0

Upgrade

Fixes Applied for Speech Recognition Issues

Date: November 5, 2024 Issues Fixed: Language code errors and training data collator problems

Problems Identified

1. Language Code Issue

  • Error: Unsupported language: tw
  • Root Cause: Whisper doesn't officially support "tw" as a language code
  • Impact: Server returning 500 errors, transcription failing

2. Data Collator Issue

  • Error: 'list' object has no attribute 'keys'
  • Root Cause: Feature extractor expected dictionary format but received list
  • Impact: Training pipeline failing during data loading

Fixes Applied

1. Language Code Fixes

A. Updated Configuration (config/config.py)

# Before:
"language": "tw",  # Twi language code for fine-tuned model

# After:
"language": None,  # Auto-detect language (Whisper doesn't officially support 'tw')

B. Updated Speech Recognizer (src/speech_recognizer.py)

  • Changed default language parameter from "tw" to None
  • Updated transcription methods to use auto-detection
  • Modified both custom and pre-trained model paths

C. Updated API Server (src/api_server.py)

  • Changed default language from "tw" to None
  • Updated API endpoint documentation
  • Modified request models to use auto-detection

D. Updated Training Script (train_whisper_twi.py)

  • Changed language configuration to use auto-detection
  • Ensures training doesn't fail on unsupported language codes

2. Data Collator Fix

Fixed in train_whisper_twi.py

# Before (line 234):
batch = self.processor.feature_extractor.pad(
    input_features, return_tensors="pt", padding=True
)

# After:
batch = self.processor.feature_extractor.pad(
    {"input_features": input_features}, return_tensors="pt", padding=True
)

Explanation: The feature extractor's pad method expects a dictionary with an "input_features" key, not a raw list of features.

Files Modified

  1. config/config.py - Updated language configuration
  2. src/speech_recognizer.py - Fixed language defaults and transcription methods
  3. src/api_server.py - Updated API endpoints and request models
  4. train_whisper_twi.py - Fixed data collator and language configuration
  5. test_fixes.py - Created test script to verify fixes

Testing

A comprehensive test script (test_fixes.py) has been created to verify:

  • Configuration loads correctly with new language settings
  • Data collator works with proper dictionary format
  • API server initializes without language errors
  • Speech recognizer handles auto-detection properly

How Whisper Will Handle Twi

Since Whisper doesn't officially support "tw" language code:

  1. Auto-Detection: Whisper will attempt to detect the language automatically
  2. Similar Languages: May detect as related languages (e.g., "ak" for Akan family)
  3. Fine-Tuned Model: Once trained on Twi data, the model will better handle Twi regardless of language code
  4. Fallback: If detection fails, Whisper defaults to English but transcription may still work

Next Steps

1. Test the Fixes

cd optimized_engine
python test_fixes.py

2. Run Quick Training

python run_training.py --quick

3. Start Server (if training succeeds)

python main.py server

4. Test Endpoints

# Health check
curl http://localhost:8000/health

# Test transcription (replace with actual audio file)
curl -X POST "http://localhost:8000/test-intent" \
  -F "file=@path/to/audio.wav"

Expected Behavior After Fixes

  1. Server Startup: Should start without language errors
  2. Transcription: Will use auto-detection instead of failing on "tw"
  3. Training: Data collator will properly format input features
  4. API Responses: Will include detected language instead of hardcoded "tw"

Monitoring Language Detection

The system will now:

  • Log detected languages for monitoring
  • Allow manual language specification if needed
  • Fall back gracefully when detection is uncertain
  • Provide confidence scores for transcriptions

Additional Recommendations

  1. Data Collection: Monitor what languages Whisper detects for your Twi audio
  2. Fine-Tuning: The custom model training will improve Twi recognition regardless of language codes
  3. Validation: Test with various Twi dialects to ensure broad coverage
  4. Fallbacks: Consider post-processing if Whisper's auto-detection needs correction

Troubleshooting

If issues persist:

  1. Check Logs: Look for language-related warnings in server logs
  2. Audio Quality: Ensure audio files are clear and properly formatted
  3. Model Loading: Verify Whisper models download correctly
  4. Dependencies: Ensure all packages are compatible versions

The fixes address the immediate blocking issues and should allow the training pipeline and server to run successfully.