Upload folder using huggingface_hub

Browse files

Files changed (6) hide show

README.md +63 -0
config.json +15 -0
merges.txt +0 -0
model.safetensors +3 -0
tokenizer_config.json +40 -0
vocab.json +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,63 @@

+---
+library_name: mlx-audio-plus
+base_model:
+- FunAudioLLM/CosyVoice2-0.5B
+tags:
+- mlx
+- tts
+- cosyvoice2
+pipeline_tag: text-to-speech
+language:
+- en
+- zh
+- ja
+- ko
+---
+# mlx-community/CosyVoice2-0.5B-8bit
+This model was converted to MLX format from [FunAudioLLM/CosyVoice2-0.5B](https://huggingface.co/FunAudioLLM/CosyVoice2-0.5B) using [mlx-audio-plus](https://github.com/DePasqualeOrg/mlx-audio-plus) version **0.1.2**.
+## Usage
+```bash
+pip install -U mlx-audio-plus
+```
+### Inference Modes
+| Mode | Parameters | Description |
+|------|------------|-------------|
+| Cross-lingual | `ref_audio` | Zero-shot TTS (default) |
+| Zero-shot | `ref_audio` + `ref_text` | Better quality with transcription |
+| Instruct | `ref_audio` + `instruct_text` | Style control (e.g., "speak slowly") |
+| Voice Conversion | `source_audio` + `ref_audio` | Convert audio to target voice |
+### Command line
+```bash
+# Cross-lingual (default)
+mlx_audio.tts --model mlx-community/CosyVoice2-0.5B-8bit --text "Hello!" --ref_audio ref.wav
+# Zero-shot (with transcription)
+mlx_audio.tts --model mlx-community/CosyVoice2-0.5B-8bit --text "Hello!" --ref_audio ref.wav --ref_text "Transcription of ref audio."
+# Instruct (style control)
+mlx_audio.tts --model mlx-community/CosyVoice2-0.5B-8bit --text "Hello!" --ref_audio ref.wav --instruct_text "Speak slowly and calmly"
+# Voice Conversion
+mlx_audio.tts --model mlx-community/CosyVoice2-0.5B-8bit --source_audio source.wav --ref_audio ref.wav
+```
+### Python
+```python
+from mlx_audio.tts.generate import generate_audio
+generate_audio(
+    text="Hello, this is CosyVoice2 on MLX!",
+    model="mlx-community/CosyVoice2-0.5B-8bit",
+    ref_audio="reference.wav",
+    file_prefix="output",
+)
+```

config.json ADDED Viewed

	@@ -0,0 +1,15 @@

+{
+  "model_type": "cosyvoice2",
+  "version": "0.5B",
+  "sample_rate": 24000,
+  "mel_channels": 80,
+  "speech_token_size": 6561,
+  "dtype": "float16",
+  "quantization": {
+    "bits": 8,
+    "group_size": 64,
+    "quantized_components": [
+      "tokenizer/model.layers"
+    ]
+  }
+}

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:94cb536647f64d1145f3031093a8caf69ba8074b46919499e4f5ae357cb13a2a
+size 957064122

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,40 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": ["<|im_start|>", "<|im_end|>"],
+  "bos_token": null,
+  "chat_template": "{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n' }}{% endif %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "model_max_length": 32768,
+  "pad_token": "<|endoftext|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff