File size: 4,251 Bytes
d0385b5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 | # Voice and language adaptation
Inflect v2 is an open-weight, inference-first release. The public checkpoint is
the complete deployable generator, not a resumable training snapshot: it does
not contain optimizer state, discriminators, training manifests, or the private
corpus-construction pipeline.
This document separates what is technically possible from what is supported.
## Current support level
| Task | Status | What that means |
| --- | --- | --- |
| Local inference | Supported | Tested through the packaged Python API and CLI. |
| Continue generator training | Research use | The released generator can initialize a compatible VITS-family training stack. |
| Replace the fixed voice | Research use | Requires a new single-speaker corpus and careful decoder adaptation. |
| Add a selectable second voice | Not supported | The release architecture has no speaker-ID or speaker-embedding interface. |
| Add a language | Not supported | Requires a new frontend, symbol inventory, embeddings, alignments, and data. |
| Zero-shot cloning | Not supported | There is no reference encoder or cloning interface. |
## What a compatible adaptation stack needs
The release config describes the deployable generator:
- 24 kHz mono waveform output;
- 16,384-sample training segments;
- English phoneme input with interspersed blanks;
- stochastic latent synthesis and monotonic alignment;
- an integrated waveform decoder.
A separately reconstructed trainer must add the training-only components that
are intentionally absent from the release:
1. waveform and period discriminators compatible with the decoder;
2. optimizer and learning-rate schedule state;
3. mel, KL, duration, feature-matching, and adversarial losses;
4. data loading, segmentation, and alignment handling;
5. validation synthesis and checkpoint export.
Loading `model.pth` as the generator initialization is reasonable. Treating it
as a complete resumable checkpoint is not.
## Data contract
Use a UTF-8 manifest with one verified transcript and one audio path per row.
The audio should be:
- mono, consistently sampled, and free of clipping;
- one speaker with stable recording conditions;
- trimmed without cutting consonants or sentence endings;
- diverse in phonemes, punctuation, sentence shape, numbers, and names;
- split into train and held-out evaluation sets before training.
Run transcript verification, duration checks, silence checks, clipping checks,
and speaker-consistency checks before training. Voice adaptation quality is
usually limited by data coverage and decoder mismatch before it is limited by
the number of optimization steps.
## Conservative warm-start sequence
This is research guidance, not a guaranteed recipe:
1. Load the public generator and freshly initialize training-only components.
2. Begin with a low generator learning rate.
3. Temporarily freeze most of the waveform decoder while the text and latent
paths adapt to the new speaker.
4. Unfreeze the decoder gradually and monitor high-frequency buzz, sibilance,
identity drift, and intelligibility.
5. Export inference-only candidates frequently.
6. Select by held-out listening and matched evaluation, not training loss.
Do not publish an adapted voice as equivalent to the release voice without a
matched unseen-text evaluation.
## Required evaluation
Compare every candidate with the frozen release on identical prompts and seeds:
- semantic WER with more than one ASR evaluator;
- clipping, silence, loudness, and ending diagnostics;
- a learned quality predictor used only as a diagnostic;
- speaker consistency for voice replacement;
- blinded human listening on unseen text.
An adaptation that improves speaker similarity while reducing intelligibility,
stability, or audio quality is not a successful adaptation.
## Why there is no one-command trainer
A nominal script would be easy to publish; a trustworthy trainer is not. The
project will call adaptation supported only after a clean environment can
produce a usable checkpoint, export it correctly, and pass the same package and
evaluation gates as the release models. Until then, the honest interface is the
tested inference package plus the architecture/configuration needed for
independent research.
|