Nemotron 3.5 Core ML encoder export experiment

Experimental 320 ms Nemotron bundle for testing the ANE specialization failure on A16 described in speech-swift issue 503. Only the encoder conversion target changes from iOS 18 to iOS 17. The export preserves all 296 published encoder palettes and copies the decoder, joint network, and runtime metadata from the published bundle without changes.

Property Value
Architecture Cache-aware FastConformer encoder and RNN-T decoder
Parameters 600 million
Format Compiled Core ML .mlmodelc
Encoder weights Published 8-bit palette indices and FP16 lookup tables
Decoder and joint weights FP16, unchanged
Runtime file size 613 MiB
Audio 16 kHz mono
Streaming chunk 320 ms
Encoder minimum deployment target iOS 17
Decoder and joint minimum deployment target iOS 18, unchanged
Baseline revision 447095fe87b480b5e6a15367f135303d479de8ac
Status Experimental; affected-iPhone validation pending

This bundle does not lower the Swift SDK's iOS 18 requirement. A successful macOS load does not establish compatibility with the affected iPhone.

Files

File Size Purpose
encoder.mlmodelc/ 565.4 MiB Encoder compiled for the iOS 17 operation set
decoder.mlmodelc/ 28.5 MiB Unchanged published RNN-T decoder
joint.mlmodelc/ 18.0 MiB Unchanged published joint network
config.json 589 B Unchanged streaming geometry
vocab.json 230.6 KiB Unchanged vocabulary
languages.json 2.0 KiB Unchanged language map
tokenizer.model 397.0 KiB Unchanged SentencePiece tokenizer
experiment.json Small JSON; exact inventory included Source revisions, export versions, and artifact hashes
TESTING.md 6.5 KiB Published versus candidate A/B instructions
testing/LoadProbe.swift 7.8 KiB Standalone iOS/macOS Core ML load probe
testing/test_audio.wav 625.0 KiB Real-speech smoke-test fixture, 20 s, mono 16 kHz
testing/mac-validation.json 30.9 KiB Local load, encoder-output, and SDK test results
testing/fixture-provenance.json <1 KiB Original fixture and resampling hashes

Validation

M5 Pro, macOS 26.6.2 (25G83):

  • All 296 encoder palette lookup tables and index arrays match the published values byte for byte. The decoder, joint, and runtime metadata are unchanged.
  • All six encoder outputs match the published CPU encoder exactly across 67 calls with real speech, silence, partial chunks, and full streaming caches.
  • Four SDK tests passed in each of three separate processes: published CPU+ANE, candidate CPU+ANE, and candidate CPU-only. Batch, streaming, and word-boosted transcripts matched the published result. The boosting engagement test also passed in all three configurations.
  • The supplied probe compiles with Swift 6 on macOS and typechecks for arm64 iOS 18. The export-helper suite passed all 14 unit tests.

The SDK tests used speech-swift revision 7fc8f6c2b7847cad17641cf294b2854e20d936a8. Validation covers one English fixture; it does not establish multilingual accuracy or A16 compatibility.

Local load and encoder timing

The standalone Core ML probe ran one bundle and compute configuration per process. Prediction measurements use synthetic input, not full ASR.

Bundle / compute units First observed encoder load Same-process reload Median repeated encoder prediction
Published / CPU+ANE 10.53 s 76.5 ms 8.09 ms
Candidate / CPU+ANE, first process 6.63 s 62.7 ms 9.10 ms
Candidate / CPU+ANE, second process 130.0 ms 65.4 ms 8.45 ms
Candidate / CPU-only 2.77 s 53.1 ms 15.50 ms

Core ML system-cache state was uncontrolled, so first-load times are observations rather than a speedup claim. CPU+ANE compute plans for both encoders prefer the ANE for 1,592 operations and CPU for 86. Planned placement does not prove execution on the ANE. Preserve device compiler logs alongside the reports.

See testing/mac-validation.json for the measurements. Repeated loads and real-speech tests on the affected iPhone remain the acceptance check.

Usage

Download the snapshot into a separate directory and follow TESTING.md. On macOS, the Python Core ML loader can check the candidate encoder directly:

import coremltools as ct
from huggingface_hub import snapshot_download

bundle = snapshot_download(
    "aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test",
    local_dir="nemotron-ios17-encoder-test",
)
encoder = ct.models.CompiledMLModel(
    f"{bundle}/encoder.mlmodelc",
    compute_units=ct.ComputeUnit.CPU_AND_NE,
)

To obtain a JSON load report on macOS:

swiftc -O -parse-as-library -D LOAD_PROBE_CLI nemotron-ios17-encoder-test/testing/LoadProbe.swift -o /tmp/nemotron-load-probe
/tmp/nemotron-load-probe nemotron-ios17-encoder-test ane candidate-ane.json

With speech-swift v0.0.27 or later, use the unchanged local-bundle API:

import CoreML
import NemotronStreamingASR

let model = try await NemotronStreamingASRModel.fromLocal(
    bundleDir: candidateBundleURL,
    computeUnits: .cpuAndNeuralEngine)

Preserve the stock bundle for the baseline comparison. The model remains an experiment until the affected iPhone passes repeated-load and real-speech batch/streaming checks.

Source and license

Upstream: NVIDIA Nemotron 3.5 ASR Streaming 0.6B, revision f3d333391852ba876df169dcc9ba902d25b6ab0b.

Baseline: published Core ML INT8 bundle. The model uses the upstream OpenMDW 1.1 license.

Links

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test

Finetuned
(60)
this model

Collection including aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test