File size: 10,840 Bytes
2be105e 1b15e82 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | Here's how you can run it on Windows 11 to convert text to speech:
**Prerequisites:**
1. **Python:** Ensure you have Python installed. You can download it from [python.org](https://www.python.org/downloads/). Version 3.8+ is recommended. During installation, make sure to check "Add Python to PATH".
2. **PyTorch:** This is the core deep learning library.
* Go to [pytorch.org](https://pytorch.org/get-started/locally/).
* Select the appropriate options for your system (PyTorch Build: Stable, Your OS: Windows, Package: Pip, Language: Python, Compute Platform: CUDA if you have an NVIDIA GPU and want to use it, otherwise CPU).
* Copy the generated `pip install` command and run it in your command prompt or terminal.
* Example for CPU: `pip3 install torch torchvision torchaudio`
* Example for CUDA 11.8: `pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118` (Check the PyTorch website for the latest CUDA version command if you have a compatible NVIDIA GPU).
3. **Transformers Library:** From Hugging Face.
```bash
pip install transformers
```
4. **SentencePiece:** Often required by tokenizers.
```bash
pip install sentencepiece
```
5. **SoundFile:** To save the audio as a `.wav` file.
```bash
pip install soundfile
```6. **Accelerate (Recommended):** For efficient model loading and execution.
```bash
pip install accelerate
```
7. **Unzip your model:** Unzip `merged_16bit_model.zip`. You should have a folder (let's assume it's named `my_sesame_tts_model` after you rename the unzipped "model" folder, or you can use "model" directly) containing files like `pytorch_model.bin` (or `.safetensors`), `config.json`, `preprocessor_config.json`, etc.
**Steps to Run Text-to-Speech:**
1. **Create a Project Folder:**
Create a new folder for your project, for example, `C:\my_tts_project`.
2. **Place Your Model:**
Move or copy the unzipped model folder (e.g., `my_sesame_tts_model`) into your project folder (`C:\my_tts_project\my_sesame_tts_model`).
3. **Create a Python Script:**
Inside your project folder (`C:\my_tts_project`), create a new Python file, for example, `run_tts.py`. Paste the following code into it:
```python
import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
import soundfile as sf
import os
# --- Configuration ---
MODEL_PATH = "./my_sesame_tts_model" # Path to your downloaded and unzipped model folder
OUTPUT_FILENAME = "output_audio.wav"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
# Set a default sampling rate if not found in config, check your model's training details
DEFAULT_SAMPLING_RATE = 24000
def text_to_speech(text, speaker_id, model_path, output_filename):
print(f"Using device: {DEVICE}")
# 1. Load processor and model
try:
print(f"Loading processor from: {model_path}")
processor = AutoProcessor.from_pretrained(model_path)
print(f"Loading model from: {model_path}")
# For merged 16-bit models, ensure your PyTorch can handle bfloat16 or float16
# If you saved as full float32, it will also work.
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_path).to(DEVICE)
model.eval() # Set model to evaluation mode
except Exception as e:
print(f"Error loading model or processor: {e}")
print("Make sure the MODEL_PATH is correct and all necessary files are present.")
return
# 2. Prepare input text
# The format "[speaker_id]text" was used in the Colab notebook.
# Adjust if your model expects a different format.
formatted_text = f"[{speaker_id}]{text}"
print(f"Input text: \"{formatted_text}\"")
try:
inputs = processor(text=formatted_text, return_tensors="pt").to(DEVICE)
except Exception as e:
print(f"Error during text processing: {e}")
return
# 3. Generate audio
print("Generating audio...")
try:
with torch.no_grad(): # Disable gradient calculations for inference
# The generate function might have specific arguments for your model
# The `output_audio=True` argument was in the Colab notebook,
# but standard AutoModelForSpeechSeq2Seq might just return audio in `output.audio`
# or directly as the output if it's the primary modality.
# Let's try a common approach first.
generated_ids = model.generate(**inputs)
# The output structure can vary.
# For Bark-like models, it might be in generated_ids["audio"] or similar.
# For Sesame models, it's often directly in the output or a specific key.
# The notebook used: audio_values = model.generate(**inputs, output_audio=True)
# Let's try to replicate that if `output_audio` is a valid generate param for this model.
# Otherwise, inspect `generated_ids`.
# For many speech models, the output is directly the audio waveforms.
if hasattr(generated_ids, 'audio'): # Check for an 'audio' attribute
audio_values = generated_ids.audio[0]
elif isinstance(generated_ids, torch.Tensor) and generated_ids.ndim >= 2 : # If output is a tensor of waveforms
audio_values = generated_ids[0]
else:
# If unsure, try to force output_audio if the model supports it (experimental)
try:
output = model.generate(**inputs, output_audio=True) # Specific to some models
audio_values = output[0] # Assuming it returns a tuple/list
except TypeError:
print("Model does not support `output_audio=True` directly or output structure is unexpected.")
print(f"Inspect `generated_ids` structure: {type(generated_ids)}")
if isinstance(generated_ids, dict): print(f"Keys: {generated_ids.keys()}")
return
except Exception as e:
print(f"Error during audio generation: {e}")
return
# 4. Post-process and save
audio_np = audio_values.cpu().to(torch.float32).numpy()
# Determine sampling rate
sampling_rate = DEFAULT_SAMPLING_RATE
if hasattr(model.config, 'sampling_rate'):
sampling_rate = model.config.sampling_rate
elif hasattr(processor, 'feature_extractor') and hasattr(processor.feature_extractor, 'sampling_rate'):
sampling_rate = processor.feature_extractor.sampling_rate
else:
print(f"Could not automatically determine sampling rate. Using default: {DEFAULT_SAMPLING_RATE} Hz.")
print(f"Using sampling rate: {sampling_rate} Hz")
try:
sf.write(output_filename, audio_np, samplerate=sampling_rate)
print(f"Audio saved to {os.path.abspath(output_filename)}")
except Exception as e:
print(f"Error saving audio file: {e}")
if __name__ == "__main__":
input_text = "Hello, this is a test of my fine-tuned text to speech model."
speaker_id_to_use = 1 # Change this to a valid speaker_id for your model
# Ensure the model path is correct relative to the script location
# or provide an absolute path.
actual_model_path = os.path.join(os.path.dirname(__file__), MODEL_PATH)
text_to_speech(input_text, speaker_id_to_use, actual_model_path, OUTPUT_FILENAME)
# Example with a different speaker
# input_text_2 = "Unsloth makes training models faster."
# speaker_id_2 = 2
# text_to_speech(input_text_2, speaker_id_2, actual_model_path, "output_audio_speaker2.wav")
```
4. **Customize the Script:**
* **`MODEL_PATH`**: Double-check this path. If `run_tts.py` is in `C:\my_tts_project` and your model is in `C:\my_tts_project\my_sesame_tts_model`, then `./my_sesame_tts_model` is correct.
* **`input_text`**: Change this to the text you want to convert.
* **`speaker_id_to_use`**: **Crucially, change this to a valid speaker ID** that your model was trained on or supports. The Colab notebook used `1` and `2` as examples.
5. **Run the Script:**
* Open a Command Prompt or PowerShell.
* Navigate to your project folder: `cd C:\my_tts_project`
* Run the script: `python run_tts.py`
This will generate an audio file (e.g., `output_audio.wav`) in your project folder.
**Important Considerations for Sesame TTS Models:**
* **Processor and Input Format:** The way you format the input text (`f"[{speaker_id}]{text}"`) is critical and depends on how the `processor` for your specific Sesame model variant was trained to expect it. The format used in the script is based on common patterns seen in the Unsloth examples. If it doesn't work, you might need to inspect the `processor.apply_chat_template` method or how inputs were prepared in the original notebook more closely.
* **Sampling Rate:** The script tries to get the sampling rate from `model.config.sampling_rate` or `processor.feature_extractor.sampling_rate`. The Colab notebook mentioned `24000` Hz, so I've set it as a default. Ensure this matches your model.
* **`model.generate()` output:** The structure of the output from `model.generate()` can vary. The provided script tries a common way and then a fallback. If audio generation fails or the output is not as expected, you might need to print `type(generated_ids)` and `generated_ids.keys()` (if it's a dict) to understand its structure and extract the audio data correctly. The Colab notebook used `audio_values = model.generate(**inputs, output_audio=True)`. The script attempts to replicate this logic if standard methods don't apply.
* **GPU Memory:** 1B parameter models, even at 16-bit, can require a decent amount of VRAM if run on GPU. If you run into CUDA out-of-memory errors, try with `DEVICE = "cpu"` (though it will be much slower).
* **Unsloth specific loading:** While `save_pretrained_merged` aims for standard Hugging Face format, if `AutoModelForSpeechSeq2Seq.from_pretrained` fails in an Unsloth-specific way, you *might* theoretically need to see if Unsloth provides a specific loading function for merged models for inference. However, this is usually not the case; `AutoModel` should work.
This comprehensive guide should get you started. The key is ensuring the paths are correct, all libraries are installed, and the input format matches what your specific fine-tuned Sesame model expects.
|