Onsei-iOS-Models / README.md
raratu's picture
Add model card, licenses, and download manifest
fdf9d2d verified
|
Raw
History Blame Contribute Delete
2.54 kB
---
license: other
language:
- ja
pipeline_tag: text-to-speech
tags:
- onnx
- ios
- on-device
- irodori-tts
---
# Onsei iOS ONNX Model Packs
ONNX model packs exported for the Onsei iOS client from the following upstream models:
- [`Aratako/Irodori-TTS-500M-v3`](https://huggingface.co/Aratako/Irodori-TTS-500M-v3), source revision `236c1e56591279fc24e3c1bf6609fc06e48dde28`
- [`Aratako/Irodori-TTS-600M-v3-VoiceDesign`](https://huggingface.co/Aratako/Irodori-TTS-600M-v3-VoiceDesign), source revision `e863a3a93e652e09afeff3e84823a206a0a60314`
- [`Aratako/Semantic-DACVAE-Japanese-32dim`](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese-32dim)
- Tokenizer from [`llm-jp/llm-jp-3-150m`](https://huggingface.co/llm-jp/llm-jp-3-150m)
These files are converted runtime artifacts, not new model weights. They are not affiliated with or endorsed by the upstream authors.
## Packs
- `500m-v3/`: standard Japanese TTS and reference-audio voice cloning.
- `600m-v3-vd/`: VoiceDesign model with caption conditioning.
Each pack contains the split `context_kv.onnx` + `dit_step.onnx` execution path used by the Onsei iOS app, DACVAE encode/decode, speaker and text encoders, duration prediction, configuration, and tokenizer data. `manifest.json` records file sizes and SHA-256 digests used by the app to verify downloads.
## Licenses and attribution
The Irodori-TTS model weights, Semantic-DACVAE model weights, and conversion code are distributed under the MIT License. The llm-jp tokenizer is distributed under Apache License 2.0. See `LICENSES/` and `THIRD_PARTY_NOTICES.md`. Each component remains subject to its own license; no relicensing is implied.
## Responsible use
The upstream model cards request that users:
1. Do not clone or impersonate a person's voice without explicit consent.
2. Do not create misleading deepfakes or synthetic speech intended to spread misinformation.
3. Recognize that text-only output may coincidentally resemble a real voice.
4. Take responsibility for compliance with applicable law.
The ONNX export used here does not add a SilentCipher watermark. Applications using these files should clearly disclose synthetic audio where appropriate and must enforce consent for voice cloning.
## Citation
```bibtex
@misc{irodori-tts-v3,
author = {Chihiro Arata},
title = {Irodori-TTS: A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control},
year = {2026},
publisher = {Hugging Face},
howpublished = {https://huggingface.co/Aratako/Irodori-TTS-500M-v3}
}
```