Jarbas's picture
Upload README.md with huggingface_hub
e040ee2 verified
|
Raw
History Blame Contribute Delete
3.65 kB
---
license: mit
license_link: https://huggingface.co/microsoft/Phi-4-multimodal-instruct/resolve/main/LICENSE
base_model:
- microsoft/Phi-4-multimodal-instruct
tags:
- automatic-speech-recognition
- onnx
- onnx-asr
- speech-llm
- phi-4
language:
- en
- zh
- de
- fr
- it
- ja
- es
- pt
library_name: onnx-asr
pipeline_tag: automatic-speech-recognition
---
# Phi-4-multimodal-instruct, audio-to-text carve-out, ONNX
This repository holds **only the speech transcription path** of
[`microsoft/Phi-4-multimodal-instruct`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct),
exported to ONNX for [onnx-asr](https://github.com/TigreGotico/onnx-asr).
The source model is a full multimodal LLM: a conformer audio encoder and a
SigLIP vision tower share one Phi-4-mini language model, which is specialised
at run time by a vision LoRA adapter or a speech LoRA adapter. This export
takes the audio branch only:
* the **speech LoRA adapter is merged** into the language model, proven
character-identical to the adapter path on four FLEURS clips before export;
* the **vision tower, the vision LoRA adapter and the vision audio projector
are dropped entirely**.
This is an audio-only carve-out. It transcribes speech. It cannot see images,
answer questions, translate or hold a conversation.
Original model by **Microsoft**, MIT licence. This export keeps that licence.
See [`REPORT.md`](REPORT.md) for the full export record, the merge proof, the
numerical error bounds and the deviations.
## Usage
The model loads on the `feat/speech-llm-qwen3-asr` branch of the TigreGotico
onnx-asr fork. **No runtime changes were needed**: the export is entirely
config-driven.
```bash
pip install "onnx-asr @ git+https://github.com/TigreGotico/onnx-asr@feat/speech-llm-qwen3-asr"
```
```python
import onnx_asr
model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx")
print(model.recognize("clip.wav"))
# roughly 3x faster and 4x smaller
model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx", quantization="int8")
```
## Languages
Phi-4-multimodal supports eight languages for audio: English, Chinese, German,
French, Italian, Japanese, Spanish and Portuguese. Other languages are out of
domain.
## Graphs
| graph | inputs | outputs |
| --- | --- | --- |
| `encoder.onnx` | `input_features` `(1, N)` f32, raw 16 kHz waveform | `audio_embeds` `(1, L, 3072)` f32 |
| `embed_tokens.onnx` | `input_ids` `(1, S)` i64 | `inputs_embeds` `(1, S, 3072)` f32 |
| `decoder.onnx` | `inputs_embeds`, `attn_bias`, `position_ids`, 32 KV pairs | `logits`, 32 KV pairs |
The Phi-4-multimodal feature extractor is the nonstandard SpeechLib
log-filterbank, so it is **baked into the encoder graph** and the model declares
`"preprocessor": "identity"`. The runtime passes the raw waveform straight
through.
## Two things to know
**Roughly 40 s of audio per call.** The upstream conformer switches to a
chunked attention path above 500 post-CNN frames, and that branch cannot be
expressed in a single ONNX graph. The exported encoder reproduces the unchunked
path exactly, which is upstream's own stated design range. Use VAD segmentation
for longer audio.
**int8 excludes the encoder.** Quantizing the conformer breaks transcription
badly, so `encoder_int8.onnx` is a full-precision copy of `encoder.onnx`. Only
`embed_tokens` and `decoder` are actually quantized. The graph-swap evidence is
in `REPORT.md`.
## Parity
ONNX fp32 is character-identical to native transformers on all four FLEURS
clips (two English, two Portuguese). The int8 variant is faithful, with small
punctuation and article differences on three of the four clips.