--- license: apache-2.0 pipeline_tag: text-to-speech tags: - zipvoice - tts - speech-synthesis - robotics - IB-Robot - ascend - onnx - edge-deployment --- # Model Card for ZipVoice-Distill (IB-Robot) ZipVoice-Distill text-to-speech (flow-matching decoder + Vocos vocoder) packaged for the [IB-Robot](https://atomgit.com/openeuler/IB_Robot) framework with two deployments sharing one contract. ## Deployments | deployment | backend | artifacts | notes | |---|---|---|---| | `ascend_310p` | Ascend ACL (Ascend310P1) | text_encoder + flow_decoder OM | fixed buckets (t256 / f1537), CFG-distilled 4-step | | `ascend_310b` | Ascend ACL (Ascend310B1) | text_encoder + flow_decoder OM (`*_310b_origin.om`) | 310B board-side synthesis | | `ubuntu_onnx` | ONNX Runtime CPU (torch profile) | dynamic-shape ONNX pair | host-side synthesis, 8-step default | Shared contract: `tensor_model/zipvoice/synthesize` (`tts.text/prompt_audio/prompt_sample_rate/prompt_text -> tts.audio`), preprocessing `emilia-zh-cn2an-jieba-pypinyin-fixed-golden-prompt-v1`, output `mono-float32-pcm-24000hz-vocos-cpu`. ## Provenance note - ONNX weights re-exported from the official k2-fsa/ZipVoice ModelScope release (`zipvoice_distill`), sha256 pinned in the manifest artifacts. - The 310P OM pair was converted from the same model with fixed shapes (`text_capacity=256`, `flow_frames=1537`, original logical revision `zipvoice-distill-310p1-bucket-2026-08-03`). - Shared assets (`tokens.txt`, `vocos/pytorch_model.bin`, `prompts/default.npz`) are byte-identical across both deployments. - Inference hyperparameters differ by design: 310P uses 4 diffusion steps, the host ONNX path defaults to 8 (see `assets/zipvoice_310p.json` vs `assets/zipvoice_onnx.json`). ## Repository Structure - `inference_manifest.json` — deployment routing (schema v3, iterative request contract) - `assets/zipvoice_310p.json` / `assets/zipvoice_onnx.json` — per-deployment runtime config - `assets/tokens.txt`, `assets/vocos/`, `assets/prompts/` — shared assets - `artifacts/ascend/ascend_310p/*.om` — compiled fixed-shape modules (310P) - `artifacts/ascend_310b/*.om` — compiled fixed-shape modules (310B) - `artifacts/onnx/*.onnx` — dynamic-shape ONNX exports ## Usage Host: `voice_tts_service` with `bundle_path: models/zipvoice`, `deployment: ubuntu_onnx`. Board: select `ascend_310p` through the unified inference runtime. ## License Code and packaging: Apache-2.0. ZipVoice model weights follow the k2-fsa/ZipVoice upstream license.