| --- |
| license: apache-2.0 |
| tags: |
| - executorch |
| - xnnpack |
| - pte |
| - on-device |
| - text-to-speech |
| base_model: |
| - hexgrad/Kokoro-82M |
| --- |
| # Kokoro-82M — ExecuTorch (text to speech, 54 voices) |
|
|
| Phonemes in, a waveform out, in one `.pte` with two methods. StyleTTS2-shaped: a |
| 12-layer phoneme BERT and a duration predictor decide how long each sound lasts, and |
| an iSTFTNet vocoder turns the stretched features into 24 kHz audio. A single |
| 256-dimensional style vector picks the voice and colours the prosody. |
|
|
| ``` |
| predict input_ids (1, N) int64, ref_s (1, 256) fp32, speed (1) fp32 |
| -> d (1, 640, N), t_en (1, 512, N), duration (N) int64 |
| vocode d, t_en, aln (1, N, F) fp32, ref_s |
| -> waveform (F * 600) fp32 @ 24 kHz |
| ``` |
|
|
| - **File**: `kokoro_82m_xnnpack_fp32.pte` — **325.4 MB**, two methods |
| - **Source**: [hexgrad/Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) — 81.8M parameters |
| - **License**: apache-2.0 |
| - **Voices**: 54, shipped separately as `voices/*.pt` in the source repo |
| - **Languages**: the graph is language-blind — it takes phoneme ids. **Verified here in |
| English and Japanese**; the voice pack also carries Spanish, French, Hindi, Italian, |
| Portuguese and Chinese voices, which are not verified on this card. |
|
|
| Both axes — phonemes and frames — are **dynamic**, and both are **exact**. Nothing is |
| padded, nothing is stretched, and there is no ladder of fixed-size methods. That took |
| some doing; see below. |
|
|
| ## Running it |
|
|
| **1. Text to phonemes, outside the graph.** Kokoro's vocabulary is IPA, and getting |
| there is [`misaki`](https://github.com/hexgrad/misaki) (or espeak-ng) — a lexicon and a |
| G2P model, not arithmetic. Same line this shelf takes with E5's prefix and Whisper's |
| mel: the recipe is here, the graph takes what the recipe produces. |
|
|
| ```python |
| from misaki import en |
| ps, _ = en.G2P(trf=False, british=False)(text) |
| ids = [vocab[c] for c in ps if c in vocab] # config.json's 178-entry vocab |
| input_ids = torch.LongTensor([[0, *ids, 0]]) # wrapped in the boundary token |
| ``` |
|
|
| For Japanese it is `misaki.ja.JAG2P()` and a `j*` voice, and nothing else changes — the |
| ids go into the same graph. Install `misaki[ja]` **with `unidic-lite`**: `unidic`'s |
| dictionary is a separate 250 MB download, and fugashi prefers `unidic` when both are |
| present, which leaves you with an empty dicdir and an error about `mecabrc`. |
|
|
| **2. Pick the style row by phoneme count.** The voice pack is 510 rows and the row is |
| `pack[len(ps) - 1]` — indexed by the length of the phoneme **string**, which is what |
| upstream's own pipeline does, and not by the number of ids, which is smaller whenever |
| the string holds a character the vocabulary does not have. |
|
|
| **3. Run `predict`,** which returns the features and one duration per phoneme. |
|
|
| **4. Build the alignment matrix yourself, at exactly `sum(duration)` frames.** |
| Upstream builds it with `repeat_interleave`, whose width is the sum of the durations — |
| the model's own output deciding the shape of its next input, which no single graph can |
| express. That is the only reason this is two methods rather than one. The same matrix |
| is comparisons only: |
|
|
| ```python |
| ends = torch.cumsum(duration, 0) |
| starts = ends - duration |
| frame = torch.arange(int(ends[-1])) # exactly sum(duration) |
| aln = ((frame[None, :] >= starts[:, None]) & |
| (frame[None, :] < ends[:, None])).float()[None] |
| ``` |
|
|
| **Give it exactly `sum(duration)` frames.** Not more — see the next section. |
|
|
| **5. Run `vocode`.** Out comes `600 * F` samples at 24 kHz. `speed` above 1 speaks |
| faster; it divides the durations before they are rounded. |
|
|
| ## Do not pad either axis |
|
|
| Both axes are dynamic, so there is no window to pad into — but it is worth saying why |
| the file is built that way, because the obvious fixed-window design does not work here |
| and the damage does not show up in a transcript. |
|
|
| | axis | what forbids padding | measured | |
| |---|---|---| |
| | phonemes | five bidirectional LSTMs — state flows in from the padding | speaking rate moves up to **19%** | |
| | frames | a bidirectional LSTM **and** `AdaIN1d`, which is `InstanceNorm` over time | log-mel **0.18–0.86** against a 0.04 noise floor | |
|
|
| The frame axis is the surprising one. `AdaIN1d` normalises over **time**, so one extra |
| frame changes the statistics the entire signal is divided by. Padding to the next |
| 16-frame rung, appending 256 frames, and padding out to 1024 all land far outside what |
| the model does to itself, and it is not a level change — taking out one global gain |
| factor leaves the distance where it was. |
|
|
| Padding with spaces rather than zeros roughly halves the damage on the phoneme axis, |
| and a recogniser transcribes **every** padded arm correctly. That is exactly why the |
| gate here is not a recogniser alone. |
|
|
| ## The LSTMs are rolled, not unrolled |
|
|
| `nn.LSTM` will not export with a dynamic sequence axis: `torch.export` pins it to |
| whatever it was traced at. The reason is that `to_edge` **unrolls** the recurrence — |
| this model's `predict` graph is 1238 ATen nodes at any length, and |
| `1651 + 108 per phoneme` in edge dialect. |
|
|
| That makes a ladder of fixed-length methods look like the only option, and then makes |
| the ladder impossible. The XNNPACK partitioner is superlinear in node count and cuts an |
| unrolled LSTM into hundreds of tiny delegates — 383 partitions at 32 phonemes — so |
| lowering one method costs: |
|
|
| | phonemes | edge nodes | lowering | |
| |---|---|---| |
| | 8 | 1651 | 33 s | |
| | 16 | 2515 | 63 s | |
| | 32 | 4243 | 153 s | |
| | 128 | 14611 | ~19 min, extrapolated | |
|
|
| A rung per phoneme count from 8 to 128 is upwards of **16 hours**, and the frame axis |
| would need its own ladder on top of that. |
|
|
| A `scan` higher-order op keeps the loop rolled. ExecuTorch lowers it, the runtime runs |
| it, and the sequence axis stays dynamic. On this model's own LSTM shape — 640 in, 256 |
| hidden each way, 128 steps: |
|
|
| | | build | edge nodes | delegates | 128 steps | |
| |---|---|---|---|---| |
| | `nn.LSTM`, one fixed length | 59.2 s | 3366 | 131 | 5.72 ms | |
| | rolled `scan`, any length | **3.1 s** | **49** | 3 | 16.65 ms | |
|
|
| Three times the runtime for one LSTM, against a build that finishes and a file that |
| takes any length. Kokoro has six of them. The whole file now builds in about two |
| minutes. |
|
|
| ## Verification |
|
|
| Ten sentences through `misaki` — five English on `af_heart`, five Japanese on |
| `jf_alpha` — each one synthesised by the `.pte` and by the unmodified upstream model in |
| eager. Three gates, because no single one works here: |
|
|
| **Durations must match exactly** — they are integers, they decide the rhythm, and they |
| come out of the half of the model that has no noise in it. **10 of 10 exact, both |
| languages.** |
|
|
| **Waveform correlation is not usable.** The vocoder's excitation carries Gaussian noise |
| and a random initial phase, so the eager model does not reproduce itself: two runs of |
| the same input correlate 0.9948. A correlation gate here measures the noise. |
|
|
| **Log-mel distance against the model's own floor** is the gate that works. Measure the |
| distance between two eager runs, then between eager and the `.pte`, and ask whether the |
| second is the first: |
|
|
| | | log-mel vs eager | eager's own floor | ratio | |
| |---|---|---|---| |
| | worst of ten | 0.0474 | 0.0455 | **1.04x** | |
| | best of ten | 0.0393 | 0.0408 | 0.96x | |
|
|
| Below and above 1.0 across the five, which is what "indistinguishable from running it |
| again" looks like — and it moves a few percent between runs, because the noise source |
| is in both arms. For scale, one extra frame of padding shows up at 4.4x, and the fp16 |
| build below at **73x**. |
|
|
| **Transcripts**, through Qwen3-ASR, scored against **eager's own transcript** rather |
| than against the sentence — what is under test is the conversion, and a recogniser |
| choosing a different kanji is not the file's doing. **CER 0.0000 on nine of ten.** |
|
|
| The tenth is worth writing down, because it is the recogniser and not the model: |
|
|
| ``` |
| この電車は東京駅に止まりますか (does this train stop at Tokyo Station) |
| eager ...東京駅に泊まりますか (stay overnight) CER 0.067 |
| pte ...東京駅に停まりますか (stop) |
| ``` |
|
|
| Same reading, different kanji — and **eager disagrees with itself here**: four runs of |
| the identical input gave 泊 once and 停 three times, and the `.pte` did the same. The |
| vocoder's noise is enough to tip a near-tie in the recogniser. Durations are identical |
| and log-mel is 1.02x the floor on this clip, so the two files are as close as eager is |
| to itself. **An ASR-only gate would have recorded this as a defect.** |
|
|
| ## Speed |
|
|
| 3.25 s of audio (50 phonemes, 130 frames) on an M-series laptop, host CPU: |
|
|
| | | ms | |
| |---|---| |
| | `predict` | 56.6 | |
| | `vocode` | 441.4 | |
| | **total** | **498.0** — 6.5x faster than real time | |
| | the same utterance in eager PyTorch | 225.5 | |
|
|
| Measured on a machine that was busy, so read these as a floor rather than a number to |
| quote. Eager is faster here and that is expected: it reaches Accelerate's LSTM and |
| convolution kernels, while about half the graph runs on portable kernels |
| (`predict` 49.6% of ops delegated, `vocode` 54.8%). No device numbers yet. |
|
|
| ## What is not in this file |
|
|
| **fp16 was built and withdrawn.** It fails three ways at once, and the first one is not |
| about ExecuTorch at all: |
|
|
| | | fp32 (this file) | fp16 (withdrawn) | |
| |---|---|---| |
| | size | 325.4 MB | 278.7 MB — a 14% saving, not the 50% the weights imply | |
| | worst log-mel / noise floor | **0.99x** | **73x** | |
| | worst CER | **0.0000** | **1.0000** — nothing intelligible | |
| | 3.25 s of audio | **498 ms** | 2104 ms | |
|
|
| **Halving this model breaks it in eager PyTorch, before any export.** At 96 phonemes the |
| vocoder returns `nan`; at 24 it survives but sits at 2.8x the noise floor. Keeping the |
| 73 data-statistic norm layers in fp32 — the usual fix for `InstanceNorm` overflow — does |
| not save it, so the overflow is not only in the norms. One duration in 96 also flips, |
| which is a rounding tie rather than a numerical failure. |
|
|
| The size is the least of it but worth knowing: the weights **do** halve to 163.4 MB, and |
| the file is still 278.7 MB, because the XNNPACK delegate carries its own copy of the |
| weights it takes. |
|
|
| **No int8 build.** Measured, not assumed: dynamic int8 quantises `nn.Linear` only, and |
| Kokoro is **66.9% Conv1d** with just **15.9%** of its weights in Linear layers, so it |
| would touch a sixth of the file. Static int8 would reach the convolutions, but a |
| vocoder's quality under quantisation has to be measured per task rather than declared, |
| and that has not been done here. |
|
|
| **The vocoder is partitioned without XNNPACK's `PermuteConfig`.** A permute inside the |
| decoder feeds two consumers outside its partition, and the delegate's output list then |
| carries that one node twice, which XNNPACK rejects with "Output node ... is already in |
| the inputs ... pass through arguments". It is the partition boundary that is wrong, not |
| the graph: the same graph lowers the moment permutes are not eligible to be one. |
|
|