Deployment
Verified release environment
The final package was validated on Windows x86-64 with Python 3.11, PyTorch 2.10, CPU inference, and CUDA inference. The runtime is ordinary PyTorch and is expected to work on modern Linux x86-64, but that combination is not yet part of the signed release matrix. Treat Python 3.11 as the reference environment rather than interpreting flexible requirements as proof of every possible version.
Install the CPU or CUDA PyTorch wheel for the target first, then install either:
python -m pip install -r requirements.txt
or the exact validation set:
python -m pip install -r requirements-tested.txt
The PyTorch build suffix remains platform-specific and is intentionally not forced by the tested requirements file.
Revision-pinned download
from huggingface_hub import snapshot_download
model_dir = snapshot_download(
"owensong/Inflect-Micro-v2",
revision="v2.0.0",
)
The v2.0.0 tag points to the frozen initial release. Use main only when you
intentionally want later documentation or runtime fixes.
Reproducibility
Keep the checkpoint, config.json, frontend files, and runtime/ directory together. Record Inflect-Micro-v2, the release revision, speed, variation, and seed with generated artifacts. Verify files against release_manifest.json when moving the package between machines.
Long text
The runtime splits long text at sentence punctuation and then at safe intra-sentence boundaries. Chunks are synthesized independently and joined with punctuation-dependent silence. This prevents unbounded inference but means a long paragraph is not one globally planned prosody pass.
Production boundaries
- Instantiate
InflectTTSonce and reuse it; loading per request wastes time. - Serialize calls to one model instance unless the surrounding application explicitly manages device memory and random seeds.
- Validate empty or untrusted input before synthesis.
- Set a fixed seed for deterministic caching and tests.
- Measure on the actual target device before claiming real-time operation.
- Apply an explicit CPU thread policy in production rather than inheriting all visible host cores.
Capacity planning
- The FP32 checkpoint is 37.53 MB for Micro and 15.97 MB for Nano.
- Peak RAM and VRAM depend on PyTorch, allocator state, input length, and device. Measure the deployed process rather than estimating memory from weight size.
- The reference runtime is non-streaming: the full chunk is returned after synthesis.
- One model instance should process one request at a time unless the application owns synchronization and random seed isolation.
Troubleshooting
- Missing phonemizer errors: reinstall
phonemizerandespeakng-loaderfromrequirements.txt. - Unexpected pronunciation: spell out ambiguous abbreviations or add context; the release frontend is English-only.
- Flat or unstable delivery: keep
variationnear the default before changing speed. - Long pauses: inspect source punctuation because chunk boundaries intentionally follow punctuation.
- Slow first request: model loading is measured separately from warm synthesis and should not happen per utterance.