--- license: apache-2.0 tags: - executorch - xnnpack - pte - on-device - text-to-speech base_model: - hexgrad/Kokoro-82M --- # Kokoro-82M — ExecuTorch (text to speech, 54 voices) Phonemes in, a waveform out, in one `.pte` with two methods. StyleTTS2-shaped: a 12-layer phoneme BERT and a duration predictor decide how long each sound lasts, and an iSTFTNet vocoder turns the stretched features into 24 kHz audio. A single 256-dimensional style vector picks the voice and colours the prosody. ``` predict input_ids (1, N) int64, ref_s (1, 256) fp32, speed (1) fp32 -> d (1, 640, N), t_en (1, 512, N), duration (N) int64 vocode d, t_en, aln (1, N, F) fp32, ref_s -> waveform (F * 600) fp32 @ 24 kHz ``` - **File**: `kokoro_82m_xnnpack_fp32.pte` — **325.4 MB**, two methods - **Source**: [hexgrad/Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) — 81.8M parameters - **License**: apache-2.0 - **Voices**: 54, shipped separately as `voices/*.pt` in the source repo - **Languages**: the graph is language-blind — it takes phoneme ids. **Verified here in English and Japanese**; the voice pack also carries Spanish, French, Hindi, Italian, Portuguese and Chinese voices, which are not verified on this card. Both axes — phonemes and frames — are **dynamic**, and both are **exact**. Nothing is padded, nothing is stretched, and there is no ladder of fixed-size methods. That took some doing; see below. ## Running it **1. Text to phonemes, outside the graph.** Kokoro's vocabulary is IPA, and getting there is [`misaki`](https://github.com/hexgrad/misaki) (or espeak-ng) — a lexicon and a G2P model, not arithmetic. Same line this shelf takes with E5's prefix and Whisper's mel: the recipe is here, the graph takes what the recipe produces. ```python from misaki import en ps, _ = en.G2P(trf=False, british=False)(text) ids = [vocab[c] for c in ps if c in vocab] # config.json's 178-entry vocab input_ids = torch.LongTensor([[0, *ids, 0]]) # wrapped in the boundary token ``` For Japanese it is `misaki.ja.JAG2P()` and a `j*` voice, and nothing else changes — the ids go into the same graph. Install `misaki[ja]` **with `unidic-lite`**: `unidic`'s dictionary is a separate 250 MB download, and fugashi prefers `unidic` when both are present, which leaves you with an empty dicdir and an error about `mecabrc`. **2. Pick the style row by phoneme count.** The voice pack is 510 rows and the row is `pack[len(ps) - 1]` — indexed by the length of the phoneme **string**, which is what upstream's own pipeline does, and not by the number of ids, which is smaller whenever the string holds a character the vocabulary does not have. **3. Run `predict`,** which returns the features and one duration per phoneme. **4. Build the alignment matrix yourself, at exactly `sum(duration)` frames.** Upstream builds it with `repeat_interleave`, whose width is the sum of the durations — the model's own output deciding the shape of its next input, which no single graph can express. That is the only reason this is two methods rather than one. The same matrix is comparisons only: ```python ends = torch.cumsum(duration, 0) starts = ends - duration frame = torch.arange(int(ends[-1])) # exactly sum(duration) aln = ((frame[None, :] >= starts[:, None]) & (frame[None, :] < ends[:, None])).float()[None] ``` **Give it exactly `sum(duration)` frames.** Not more — see the next section. **5. Run `vocode`.** Out comes `600 * F` samples at 24 kHz. `speed` above 1 speaks faster; it divides the durations before they are rounded. ## Do not pad either axis Both axes are dynamic, so there is no window to pad into — but it is worth saying why the file is built that way, because the obvious fixed-window design does not work here and the damage does not show up in a transcript. | axis | what forbids padding | measured | |---|---|---| | phonemes | five bidirectional LSTMs — state flows in from the padding | speaking rate moves up to **19%** | | frames | a bidirectional LSTM **and** `AdaIN1d`, which is `InstanceNorm` over time | log-mel **0.18–0.86** against a 0.04 noise floor | The frame axis is the surprising one. `AdaIN1d` normalises over **time**, so one extra frame changes the statistics the entire signal is divided by. Padding to the next 16-frame rung, appending 256 frames, and padding out to 1024 all land far outside what the model does to itself, and it is not a level change — taking out one global gain factor leaves the distance where it was. Padding with spaces rather than zeros roughly halves the damage on the phoneme axis, and a recogniser transcribes **every** padded arm correctly. That is exactly why the gate here is not a recogniser alone. ## The LSTMs are rolled, not unrolled `nn.LSTM` will not export with a dynamic sequence axis: `torch.export` pins it to whatever it was traced at. The reason is that `to_edge` **unrolls** the recurrence — this model's `predict` graph is 1238 ATen nodes at any length, and `1651 + 108 per phoneme` in edge dialect. That makes a ladder of fixed-length methods look like the only option, and then makes the ladder impossible. The XNNPACK partitioner is superlinear in node count and cuts an unrolled LSTM into hundreds of tiny delegates — 383 partitions at 32 phonemes — so lowering one method costs: | phonemes | edge nodes | lowering | |---|---|---| | 8 | 1651 | 33 s | | 16 | 2515 | 63 s | | 32 | 4243 | 153 s | | 128 | 14611 | ~19 min, extrapolated | A rung per phoneme count from 8 to 128 is upwards of **16 hours**, and the frame axis would need its own ladder on top of that. A `scan` higher-order op keeps the loop rolled. ExecuTorch lowers it, the runtime runs it, and the sequence axis stays dynamic. On this model's own LSTM shape — 640 in, 256 hidden each way, 128 steps: | | build | edge nodes | delegates | 128 steps | |---|---|---|---|---| | `nn.LSTM`, one fixed length | 59.2 s | 3366 | 131 | 5.72 ms | | rolled `scan`, any length | **3.1 s** | **49** | 3 | 16.65 ms | Three times the runtime for one LSTM, against a build that finishes and a file that takes any length. Kokoro has six of them. The whole file now builds in about two minutes. ## Verification Ten sentences through `misaki` — five English on `af_heart`, five Japanese on `jf_alpha` — each one synthesised by the `.pte` and by the unmodified upstream model in eager. Three gates, because no single one works here: **Durations must match exactly** — they are integers, they decide the rhythm, and they come out of the half of the model that has no noise in it. **10 of 10 exact, both languages.** **Waveform correlation is not usable.** The vocoder's excitation carries Gaussian noise and a random initial phase, so the eager model does not reproduce itself: two runs of the same input correlate 0.9948. A correlation gate here measures the noise. **Log-mel distance against the model's own floor** is the gate that works. Measure the distance between two eager runs, then between eager and the `.pte`, and ask whether the second is the first: | | log-mel vs eager | eager's own floor | ratio | |---|---|---|---| | worst of ten | 0.0474 | 0.0455 | **1.04x** | | best of ten | 0.0393 | 0.0408 | 0.96x | Below and above 1.0 across the five, which is what "indistinguishable from running it again" looks like — and it moves a few percent between runs, because the noise source is in both arms. For scale, one extra frame of padding shows up at 4.4x, and the fp16 build below at **73x**. **Transcripts**, through Qwen3-ASR, scored against **eager's own transcript** rather than against the sentence — what is under test is the conversion, and a recogniser choosing a different kanji is not the file's doing. **CER 0.0000 on nine of ten.** The tenth is worth writing down, because it is the recogniser and not the model: ``` この電車は東京駅に止まりますか (does this train stop at Tokyo Station) eager ...東京駅に泊まりますか (stay overnight) CER 0.067 pte ...東京駅に停まりますか (stop) ``` Same reading, different kanji — and **eager disagrees with itself here**: four runs of the identical input gave 泊 once and 停 three times, and the `.pte` did the same. The vocoder's noise is enough to tip a near-tie in the recogniser. Durations are identical and log-mel is 1.02x the floor on this clip, so the two files are as close as eager is to itself. **An ASR-only gate would have recorded this as a defect.** ## Speed 3.25 s of audio (50 phonemes, 130 frames) on an M-series laptop, host CPU: | | ms | |---|---| | `predict` | 56.6 | | `vocode` | 441.4 | | **total** | **498.0** — 6.5x faster than real time | | the same utterance in eager PyTorch | 225.5 | Measured on a machine that was busy, so read these as a floor rather than a number to quote. Eager is faster here and that is expected: it reaches Accelerate's LSTM and convolution kernels, while about half the graph runs on portable kernels (`predict` 49.6% of ops delegated, `vocode` 54.8%). No device numbers yet. ## What is not in this file **fp16 was built and withdrawn.** It fails three ways at once, and the first one is not about ExecuTorch at all: | | fp32 (this file) | fp16 (withdrawn) | |---|---|---| | size | 325.4 MB | 278.7 MB — a 14% saving, not the 50% the weights imply | | worst log-mel / noise floor | **0.99x** | **73x** | | worst CER | **0.0000** | **1.0000** — nothing intelligible | | 3.25 s of audio | **498 ms** | 2104 ms | **Halving this model breaks it in eager PyTorch, before any export.** At 96 phonemes the vocoder returns `nan`; at 24 it survives but sits at 2.8x the noise floor. Keeping the 73 data-statistic norm layers in fp32 — the usual fix for `InstanceNorm` overflow — does not save it, so the overflow is not only in the norms. One duration in 96 also flips, which is a rounding tie rather than a numerical failure. The size is the least of it but worth knowing: the weights **do** halve to 163.4 MB, and the file is still 278.7 MB, because the XNNPACK delegate carries its own copy of the weights it takes. **No int8 build.** Measured, not assumed: dynamic int8 quantises `nn.Linear` only, and Kokoro is **66.9% Conv1d** with just **15.9%** of its weights in Linear layers, so it would touch a sixth of the file. Static int8 would reach the convolutions, but a vocoder's quality under quantisation has to be measured per task rather than declared, and that has not been done here. **The vocoder is partitioned without XNNPACK's `PermuteConfig`.** A permute inside the decoder feeds two consumers outside its partition, and the delegate's output list then carries that one node twice, which XNNPACK rejects with "Output node ... is already in the inputs ... pass through arguments". It is the partition boundary that is wrong, not the graph: the same graph lowers the moment permutes are not eligible to be one.