mlboydaisuke's picture
Upload README.md with huggingface_hub
43e9990 verified
|
Raw
History Blame Contribute Delete
11 kB
---
license: apache-2.0
tags:
- executorch
- xnnpack
- pte
- on-device
- text-to-speech
base_model:
- hexgrad/Kokoro-82M
---
# Kokoro-82M — ExecuTorch (text to speech, 54 voices)
Phonemes in, a waveform out, in one `.pte` with two methods. StyleTTS2-shaped: a
12-layer phoneme BERT and a duration predictor decide how long each sound lasts, and
an iSTFTNet vocoder turns the stretched features into 24 kHz audio. A single
256-dimensional style vector picks the voice and colours the prosody.
```
predict input_ids (1, N) int64, ref_s (1, 256) fp32, speed (1) fp32
-> d (1, 640, N), t_en (1, 512, N), duration (N) int64
vocode d, t_en, aln (1, N, F) fp32, ref_s
-> waveform (F * 600) fp32 @ 24 kHz
```
- **File**: `kokoro_82m_xnnpack_fp32.pte`**325.4 MB**, two methods
- **Source**: [hexgrad/Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) — 81.8M parameters
- **License**: apache-2.0
- **Voices**: 54, shipped separately as `voices/*.pt` in the source repo
- **Languages**: the graph is language-blind — it takes phoneme ids. **Verified here in
English and Japanese**; the voice pack also carries Spanish, French, Hindi, Italian,
Portuguese and Chinese voices, which are not verified on this card.
Both axes — phonemes and frames — are **dynamic**, and both are **exact**. Nothing is
padded, nothing is stretched, and there is no ladder of fixed-size methods. That took
some doing; see below.
## Running it
**1. Text to phonemes, outside the graph.** Kokoro's vocabulary is IPA, and getting
there is [`misaki`](https://github.com/hexgrad/misaki) (or espeak-ng) — a lexicon and a
G2P model, not arithmetic. Same line this shelf takes with E5's prefix and Whisper's
mel: the recipe is here, the graph takes what the recipe produces.
```python
from misaki import en
ps, _ = en.G2P(trf=False, british=False)(text)
ids = [vocab[c] for c in ps if c in vocab] # config.json's 178-entry vocab
input_ids = torch.LongTensor([[0, *ids, 0]]) # wrapped in the boundary token
```
For Japanese it is `misaki.ja.JAG2P()` and a `j*` voice, and nothing else changes — the
ids go into the same graph. Install `misaki[ja]` **with `unidic-lite`**: `unidic`'s
dictionary is a separate 250 MB download, and fugashi prefers `unidic` when both are
present, which leaves you with an empty dicdir and an error about `mecabrc`.
**2. Pick the style row by phoneme count.** The voice pack is 510 rows and the row is
`pack[len(ps) - 1]` — indexed by the length of the phoneme **string**, which is what
upstream's own pipeline does, and not by the number of ids, which is smaller whenever
the string holds a character the vocabulary does not have.
**3. Run `predict`,** which returns the features and one duration per phoneme.
**4. Build the alignment matrix yourself, at exactly `sum(duration)` frames.**
Upstream builds it with `repeat_interleave`, whose width is the sum of the durations —
the model's own output deciding the shape of its next input, which no single graph can
express. That is the only reason this is two methods rather than one. The same matrix
is comparisons only:
```python
ends = torch.cumsum(duration, 0)
starts = ends - duration
frame = torch.arange(int(ends[-1])) # exactly sum(duration)
aln = ((frame[None, :] >= starts[:, None]) &
(frame[None, :] < ends[:, None])).float()[None]
```
**Give it exactly `sum(duration)` frames.** Not more — see the next section.
**5. Run `vocode`.** Out comes `600 * F` samples at 24 kHz. `speed` above 1 speaks
faster; it divides the durations before they are rounded.
## Do not pad either axis
Both axes are dynamic, so there is no window to pad into — but it is worth saying why
the file is built that way, because the obvious fixed-window design does not work here
and the damage does not show up in a transcript.
| axis | what forbids padding | measured |
|---|---|---|
| phonemes | five bidirectional LSTMs — state flows in from the padding | speaking rate moves up to **19%** |
| frames | a bidirectional LSTM **and** `AdaIN1d`, which is `InstanceNorm` over time | log-mel **0.18–0.86** against a 0.04 noise floor |
The frame axis is the surprising one. `AdaIN1d` normalises over **time**, so one extra
frame changes the statistics the entire signal is divided by. Padding to the next
16-frame rung, appending 256 frames, and padding out to 1024 all land far outside what
the model does to itself, and it is not a level change — taking out one global gain
factor leaves the distance where it was.
Padding with spaces rather than zeros roughly halves the damage on the phoneme axis,
and a recogniser transcribes **every** padded arm correctly. That is exactly why the
gate here is not a recogniser alone.
## The LSTMs are rolled, not unrolled
`nn.LSTM` will not export with a dynamic sequence axis: `torch.export` pins it to
whatever it was traced at. The reason is that `to_edge` **unrolls** the recurrence —
this model's `predict` graph is 1238 ATen nodes at any length, and
`1651 + 108 per phoneme` in edge dialect.
That makes a ladder of fixed-length methods look like the only option, and then makes
the ladder impossible. The XNNPACK partitioner is superlinear in node count and cuts an
unrolled LSTM into hundreds of tiny delegates — 383 partitions at 32 phonemes — so
lowering one method costs:
| phonemes | edge nodes | lowering |
|---|---|---|
| 8 | 1651 | 33 s |
| 16 | 2515 | 63 s |
| 32 | 4243 | 153 s |
| 128 | 14611 | ~19 min, extrapolated |
A rung per phoneme count from 8 to 128 is upwards of **16 hours**, and the frame axis
would need its own ladder on top of that.
A `scan` higher-order op keeps the loop rolled. ExecuTorch lowers it, the runtime runs
it, and the sequence axis stays dynamic. On this model's own LSTM shape — 640 in, 256
hidden each way, 128 steps:
| | build | edge nodes | delegates | 128 steps |
|---|---|---|---|---|
| `nn.LSTM`, one fixed length | 59.2 s | 3366 | 131 | 5.72 ms |
| rolled `scan`, any length | **3.1 s** | **49** | 3 | 16.65 ms |
Three times the runtime for one LSTM, against a build that finishes and a file that
takes any length. Kokoro has six of them. The whole file now builds in about two
minutes.
## Verification
Ten sentences through `misaki` — five English on `af_heart`, five Japanese on
`jf_alpha` — each one synthesised by the `.pte` and by the unmodified upstream model in
eager. Three gates, because no single one works here:
**Durations must match exactly** — they are integers, they decide the rhythm, and they
come out of the half of the model that has no noise in it. **10 of 10 exact, both
languages.**
**Waveform correlation is not usable.** The vocoder's excitation carries Gaussian noise
and a random initial phase, so the eager model does not reproduce itself: two runs of
the same input correlate 0.9948. A correlation gate here measures the noise.
**Log-mel distance against the model's own floor** is the gate that works. Measure the
distance between two eager runs, then between eager and the `.pte`, and ask whether the
second is the first:
| | log-mel vs eager | eager's own floor | ratio |
|---|---|---|---|
| worst of ten | 0.0474 | 0.0455 | **1.04x** |
| best of ten | 0.0393 | 0.0408 | 0.96x |
Below and above 1.0 across the five, which is what "indistinguishable from running it
again" looks like — and it moves a few percent between runs, because the noise source
is in both arms. For scale, one extra frame of padding shows up at 4.4x, and the fp16
build below at **73x**.
**Transcripts**, through Qwen3-ASR, scored against **eager's own transcript** rather
than against the sentence — what is under test is the conversion, and a recogniser
choosing a different kanji is not the file's doing. **CER 0.0000 on nine of ten.**
The tenth is worth writing down, because it is the recogniser and not the model:
```
この電車は東京駅に止まりますか (does this train stop at Tokyo Station)
eager ...東京駅に泊まりますか (stay overnight) CER 0.067
pte ...東京駅に停まりますか (stop)
```
Same reading, different kanji — and **eager disagrees with itself here**: four runs of
the identical input gave 泊 once and 停 three times, and the `.pte` did the same. The
vocoder's noise is enough to tip a near-tie in the recogniser. Durations are identical
and log-mel is 1.02x the floor on this clip, so the two files are as close as eager is
to itself. **An ASR-only gate would have recorded this as a defect.**
## Speed
3.25 s of audio (50 phonemes, 130 frames) on an M-series laptop, host CPU:
| | ms |
|---|---|
| `predict` | 56.6 |
| `vocode` | 441.4 |
| **total** | **498.0** — 6.5x faster than real time |
| the same utterance in eager PyTorch | 225.5 |
Measured on a machine that was busy, so read these as a floor rather than a number to
quote. Eager is faster here and that is expected: it reaches Accelerate's LSTM and
convolution kernels, while about half the graph runs on portable kernels
(`predict` 49.6% of ops delegated, `vocode` 54.8%). No device numbers yet.
## What is not in this file
**fp16 was built and withdrawn.** It fails three ways at once, and the first one is not
about ExecuTorch at all:
| | fp32 (this file) | fp16 (withdrawn) |
|---|---|---|
| size | 325.4 MB | 278.7 MB — a 14% saving, not the 50% the weights imply |
| worst log-mel / noise floor | **0.99x** | **73x** |
| worst CER | **0.0000** | **1.0000** — nothing intelligible |
| 3.25 s of audio | **498 ms** | 2104 ms |
**Halving this model breaks it in eager PyTorch, before any export.** At 96 phonemes the
vocoder returns `nan`; at 24 it survives but sits at 2.8x the noise floor. Keeping the
73 data-statistic norm layers in fp32 — the usual fix for `InstanceNorm` overflow — does
not save it, so the overflow is not only in the norms. One duration in 96 also flips,
which is a rounding tie rather than a numerical failure.
The size is the least of it but worth knowing: the weights **do** halve to 163.4 MB, and
the file is still 278.7 MB, because the XNNPACK delegate carries its own copy of the
weights it takes.
**No int8 build.** Measured, not assumed: dynamic int8 quantises `nn.Linear` only, and
Kokoro is **66.9% Conv1d** with just **15.9%** of its weights in Linear layers, so it
would touch a sixth of the file. Static int8 would reach the convolutions, but a
vocoder's quality under quantisation has to be measured per task rather than declared,
and that has not been done here.
**The vocoder is partitioned without XNNPACK's `PermuteConfig`.** A permute inside the
decoder feeds two consumers outside its partition, and the delegate's output list then
carries that one node twice, which XNNPACK rejects with "Output node ... is already in
the inputs ... pass through arguments". It is the partition boundary that is wrong, not
the graph: the same graph lowers the moment permutes are not eligible to be one.