--- language: - ar tags: - text-to-speech - tts - arabic - voice-cloning - zero-shot - audio - vibevoice license: apache-2.0 pipeline_tag: text-to-speech --- # 🎙️ Arabic_TTS by theBOrganization

High-Fidelity, Zero-Shot Arabic Text-to-Speech

Bringing natural, expressive, and culturally accurate Arabic speech to life.


## 🌟 Overview **Arabic_TTS** is a state-of-the-art Text-to-Speech model specifically optimized for the Arabic language. Built on the powerful **VibeVoice** architecture, this model goes beyond standard TTS by offering **zero-shot voice cloning**. Simply provide a short reference audio clip, and the model will synthesize your input text in Arabic while perfectly capturing the timbre, tone, and unique characteristics of the reference speaker. ### ✨ Key Features * 🗣️ **Native Arabic Fluency:** Handles complex Arabic morphology, diacritics (Tashkeel) with high fidelity. * 🎭 **Zero-Shot Voice Cloning:** Clone any voice from a single reference audio sample without fine-tuning. * 🎛️ **Controllable Generation:** Built-in Classifier-Free Guidance (CFG) and temperature sampling for expressive and diverse outputs. * ⚡ **Optimized Inference:** Supports CUDA (bfloat16), Apple Silicon MPS (float32), and CPU for maximum efficiency. --- ## 🎧 Audio Samples Listen to the quality of the model below. We compare the original reference voice against the synthetic generated output.
🎤 Original Reference Voice 🤖 Synthetic Generated Voice
--- ## 🛠️ Usage & Inference ### Prerequisites Ensure you have PyTorch installed. You will also need the `vibevoice` library. ```bash pip install torch torchaudio # Install vibevoice (adjust based on your specific package distribution) pip install vibevoice ``` ### Inference Code Below is the complete script to load the model, process the text and reference audio, and generate high-quality 24kHz speech. ```python import os import torch import torchaudio as ta from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor from vibevoice.modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference # --- Configuration --- model_path = "theBOrganization/Arabic_TTS" text = "SPEAKER 1: مرحبا، هذا مثال على توليد الصوت باللغة العربية باستخدام نموذج جديد." output_path = "output.wav" # Device selection with fallback device = "cuda" if torch.cuda.is_available() else ("mps" if torch.backends.mps.is_available() else "cpu") print(f"🚀 Loading model on {device}...") # Load Processor processor = VibeVoiceProcessor.from_pretrained(model_path) # Set dtype and attention implementation based on device if device == "mps": dtype = torch.float32 elif device == "cuda": dtype = torch.bfloat16 else: dtype = torch.float32 # Note: float32 is recommended for CPU stability attn_impl = "sdpa" # Load Model model = VibeVoiceForConditionalGenerationInference.from_pretrained( model_path, torch_dtype=dtype, attn_implementation=attn_impl, device_map=device if device != "mps" else None, ) if device == "mps": model.to("mps") model.eval() model.set_ddpm_inference_steps(10) # 10 steps for DDPM inference # --- Prepare Inputs --- reference_voice = "prompt.wav" # Replace with your reference audio path print(f"🎙️ Using reference voice: {reference_voice}") inputs = processor( text=[text], voice_samples=[[reference_voice]], # Batch of one speaker list return_tensors="pt", padding=True, ) # Move tensors to the correct device inputs = { k: v.to(device) if torch.is_tensor(v) else v for k, v in inputs.items() } # --- Generate Audio --- print("🎧 Generating audio...") with torch.no_grad(): outputs = model.generate( **inputs, cfg_scale=1.3, tokenizer=processor.tokenizer, generation_config={ 'do_sample': True, 'temperature': 0.5 }, ) # --- Save Output --- audio = outputs.speech_outputs[0].cpu().float() # shape: [1, T] or [T] if audio.ndim == 1: audio = audio.unsqueeze(0).float() sample_rate = 24000 ta.save(output_path, audio, sample_rate) print(f"✅ Successfully saved audio to {output_path}") ``` ### 📝 Input Formatting The model expects text to be prefixed with a speaker identifier. * **Format:** `SPEAKER X: ` * **Example:** `SPEAKER 1: أهلاً وسهلاً بكم في موقعنا.` --- ## ⚙️ Model Details | Feature | Specification | | :--- | :--- | | **Architecture** | VibeVoice (Diffusion / Flow-Matching based) | | **Sampling Rate** | 24,000 Hz | | **Inference Steps** | 10 (DDPM) | | **Supported Languages** | Arabic (Modern Standard) | --- ## ⚠️ Limitations & Ethical Considerations While **Arabic_TTS** produces highly realistic speech, users must adhere to responsible AI practices: 1. **Consent:** Do not use the zero-shot voice cloning feature to clone voices without the explicit consent of the speaker. 2. **Misinformation:** Do not use this model to generate deceptive audio (deepfakes) for malicious purposes, fraud, or political manipulation. 3. **Dialects:** While the model excels at Modern Standard Arabic (MSA), extreme regional dialects or heavy code-switching (mixing Arabic and English in the same sentence) may result in slight pronunciation artifacts. --- **Developed with ❤️ by [theBOrganization]** For inquiries, collaborations, or bug reports, please open an issue on the repository or contact us directly.