Parakeet TDT 0.6B v3 — Core AI (8-bit palettized)

Community conversion by spybyscript of NVIDIA's Parakeet TDT 0.6B v3 into Apple's Core AI .aimodel format for OS 27. It's exported directly from the original PyTorch weights, with no training or fine-tuning. This is not an official NVIDIA or Apple release.

The bundle in pal8/ has four component graphs: frontend, encoder, predictor and joint network. A host-side TDT greedy decoder turns them into text for recordings of up to 35 seconds, with the tokenizer, tensor contract, checksums and validation alongside. See INTEGRATION.md.

Related formats: spybyscript's Android LiteRT conversion. This Core AI export uses the original checkpoint, not the LiteRT files.

Source and license

  • Original revision: 541d1f99c6b0c3cd0b11a95167540bb8edefd82b.
  • Source safetensors SHA-256: 3a2026366188c8c68598edbbff92f8d11590a08e0ae2e6775544e7b07d6a5e11.
  • License: CC-BY-4.0, the same as the original. See LICENSE.txt. Attribution and the list of changes are in NOTICE.md.
  • The upstream card is kept in pal8/SOURCE_MODEL_CARD.md. Its benchmarks are upstream results, not measurements of this conversion.
  • Export: PyTorch 2.11.0, Transformers 5.17.0, coreai-torch 0.4.2, coreai-core 1.0.0b2, coreai-opt 0.2.1. See conversion/README.md.

What makes it small

A plain FP16 export of this model is 1.29 GB. A copy precompiled for one iPhone chip family added 1.46 GB more, for 2.76 GB. This bundle is 649,528,645 bytes (0.65 GB):

Component Weights Activations and interface Bytes
Frontend (DFT, mel, normalization) FP32 FP32 2.1 M
Encoder and projection 8-bit palettized (per-tensor k-means, 256 entries) FP16 610 M
Predictor (2-layer LSTM) FP16 FP32 in and out 23 M
Joint network FP16 FP32 in and out 10 M

There's no precompiled copy: Core AI specializes the bundle on first load. In an earlier study of a different model, palettized weights stayed compressed when compiled for an iPhone GPU, while 8-bit linear quantization expanded. This bundle's own iPhone behavior hasn't been measured yet.

Measured results

Conversion checks used 20 public LibriSpeech test-clean recordings, 10 speakers, 171.6 seconds, 488 reference words, plus a 29.2-second multi-window probe built from them. This is a small conversion test, not the full LibriSpeech benchmark or a broad accuracy claim.

Original model (PyTorch FP32) FP16 Core AI (not published) This bundle
Bundle — 2.76 GB 0.65 GB
Exact tokens against FP32 (Mac, native, 2 passes) reference 42/42 42/42
Human-reference WER (unicode-v1) 2.05% (10/488) 2.05% 2.05%
Load — 3.95 s 1.09 s
Warm RTF¹ — 0.0143 0.0108
Peak app footprint² — 282 MiB 322 MiB

¹ Processing seconds per audio second, from the second pass; lower is faster. ² Physical footprint of the benchmark process.

Both columns were measured on the same Mac in the same session: Apple M5 Max, 128 GiB, macOS 27.0, Core AI default specialization. Actual GPU, CPU or Neural Engine placement wasn't profiled.

  • Precision screening: before export, the encoder was screened in PyTorch at several weight precisions against the FP32 model (evidence/precision-screen.json).
    • FP16 and 8-bit palettized matched all 21 inputs exactly.
    • 6-bit (0.46 GB for the encoder) changed 3 of 21 and was not used.
  • Full results: pal8/validation.json, with the public transcripts. The recordings themselves aren't distributed.

On an iPhone

Measured on an iPhone 17 Pro Max (iOS 27) in the Read the Room app's silent fixture check: no microphone, with live audio-emotion analysis running alongside. The clips were public LibriSpeech 2094-142345-0051 (6.3 s) and two variants of it. Both bundles ran the same clips in the same app build (evidence/iphone-fixture.json).

FP16 with an h18p precompiled copy (2.76 GB) This bundle (0.65 GB)
Transcripts — identical on all 3 clips
Model load 14.3 s 4.4 s
First transcription (6.3 s clip) 0.72 s 2.71 s¹
Warm transcription 0.58 s 0.52 s
App memory (whole app) ~2.05 GB ~0.93 GB

¹ On first use, Core AI specializes the source assets for the device, and keeps the result until the next OS update. The FP16 bundle skipped this step with its precompiled copy for this chip family.

Limits

  • One phone so far. Other iPhones and iPads, thermals under sustained use, and compute placement haven't been measured.
  • English only, so far. Parakeet TDT 0.6B v3 transcribes 25 European languages. This conversion was checked on English recordings only.
  • FP16 on some iPhones. There are open reports of FP16 Core AI graphs giving wrong outputs under default specialization on some devices (apple/coreai-torch #115, #116). Check parity on each target device.
  • Scope. This is post-recording transcription. The encoder takes one fixed 15-second window; longer recordings go through the windowing in INTEGRATION.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spybyscript/parakeet-tdt-coreai

Finetuned
(92)
this model