| --- |
| license: other |
| license_name: funasr-model-license-1.1 |
| license_link: https://github.com/modelscope/FunASR/blob/main/MODEL_LICENSE |
| library_name: coremltools |
| pipeline_tag: automatic-speech-recognition |
| base_model: FunAudioLLM/SenseVoiceSmall |
| language: |
| - zh |
| - en |
| - yue |
| - ja |
| - ko |
| tags: |
| - coreml |
| - apple-silicon |
| - automatic-speech-recognition |
| - speech-understanding |
| - emotion-recognition |
| - audio-event-detection |
| - multilingual |
| - macos |
| --- |
| |
| # SenseVoiceSmall — Core ML for Apple Silicon |
|
|
| A ready-to-use Core ML conversion of [FunAudioLLM/SenseVoiceSmall](https://huggingface.co/FunAudioLLM/SenseVoiceSmall) for fast, private speech recognition and rich audio understanding on Apple Silicon. |
|
|
| SenseVoiceSmall performs transcription, spoken-language identification, speech-emotion recognition, and audio-event detection in one non-autoregressive model. It is integrated into [Hugging Mac](https://github.com/devilyouwei/hugging-mac) for building local macOS voice apps, services, plugins, and agents. |
|
|
| ## Capabilities |
|
|
| - Speech recognition for Chinese, English, Cantonese, Japanese, and Korean |
| - Automatic spoken-language identification |
| - Emotion recognition: happy, sad, angry, neutral, fearful, disgusted, and surprised |
| - Audio-event detection including music, applause, laughter, crying, coughing, sneezing, breathing, singing, and speech noise |
| - Optional inverse text normalization |
| - Short-form audio up to 30 seconds at 16 kHz |
|
|
| ## Core ML model |
|
|
| The bundle contains one multifunction ML Program with three fixed feature buckets. Hugging Mac automatically selects the smallest suitable function. |
|
|
| | Function | `features` input | Other inputs | `logits` output | |
| |---|---|---|---| |
| | `encoder_100` | `1 × 100 × 560` FP32 | `lengths`, `language_id`, `style_id`: `1` INT32 | `1 × 104 × 25055` FP16 | |
| | `encoder_250` | `1 × 250 × 560` FP32 | `lengths`, `language_id`, `style_id`: `1` INT32 | `1 × 254 × 25055` FP16 | |
| | `encoder_500` | `1 × 500 × 560` FP32 | `lengths`, `language_id`, `style_id`: `1` INT32 | `1 × 504 × 25055` FP16 | |
|
|
| - Architecture: non-autoregressive SANM encoder with CTC output |
| - Weight optimization: INT8 linear quantization |
| - Compute precision: FP16 |
| - Minimum deployment target: macOS 15 |
| - Package size: 238.7 MB |
| - Default function: `encoder_500` |
|
|
| The inputs are 560-dimensional acoustic features generated with the included `am.mvn` normalization statistics. Raw logits require CTC collapse, SentencePiece decoding, and parsing of rich language, emotion, and event tokens. The tokenizer is downloaded from the pinned upstream repository by Hugging Mac. |
|
|
| ## Use with Hugging Mac |
|
|
| ```bash |
| git clone https://github.com/devilyouwei/hugging-mac.git |
| cd hugging-mac |
| uv sync --all-packages |
| uv run hugging-mac-web |
| ``` |
|
|
| Open the Models page, install **SenseVoiceSmall Core ML**, and use it through Live Transcription or the `SpeechTranscription` and `SpeechUnderstanding` capabilities. |
|
|
| For feature extraction, Core ML invocation, CTC decoding, and structured rich-transcript parsing, see the [Hugging Mac SenseVoice SDK](https://github.com/devilyouwei/hugging-mac/tree/main/packages/hugging_mac_sdk/src/hugging_mac_sdk/models/sensevoice). |
|
|
| ## Provenance and integrity |
|
|
| - Upstream model: [FunAudioLLM/SenseVoiceSmall](https://huggingface.co/FunAudioLLM/SenseVoiceSmall) |
| - Upstream revision: `3847d57b6bdf2dd8875cb1508d2af43d80a16bf7` |
| - Conversion format: `sensevoice-small-coreml` |
| - Conversion metadata: included in `sensevoice-small-coreml/conversion.json` |
| - Directory SHA-256: `6ec1ee9247f7c717a285d809967bd64e210dbba4418e872c01a7a60bdedbb983` |
|
|
| ## License |
|
|
| The converted model retains the **FunASR Model Open Source License Agreement 1.1** and its attribution and usage conditions. Review the linked license before redistribution or commercial use. Hugging Mac is an independent open-source project and is not affiliated with or endorsed by FunAudioLLM, FunASR, or ModelScope. |
|
|