File size: 7,400 Bytes
e61ba9f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0b5a6ae
 
e61ba9f
 
 
 
 
 
 
 
 
 
 
c807c0c
 
 
 
 
 
 
 
0b5a6ae
 
c807c0c
 
 
 
 
 
 
0b5a6ae
 
 
 
 
 
 
 
 
5d38559
 
e61ba9f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0b5a6ae
 
 
 
 
 
 
 
 
 
 
 
 
e61ba9f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
---
license: apache-2.0
language:
- multilingual
pipeline_tag: automatic-speech-recognition
tags:
- audio
- automatic-speech-recognition
- whisper
- safetensors
---

# Whisper inference weights

Inference-oriented conversions of OpenAI Whisper checkpoints. The model architecture and learned parameters come from [OpenAI Whisper](https://github.com/openai/whisper) and the corresponding [`openai/whisper-*`](https://huggingface.co/models?search=openai%2Fwhisper) repositories.

This repository uses branches to keep each model at the conventional root path `model.safetensors`:

| Revision | Source | Parameters | Languages |
|---|---|---:|---|
| [`tiny`](https://huggingface.co/vpermilp/whisper/tree/tiny) | [`openai/whisper-tiny`](https://huggingface.co/openai/whisper-tiny) | 39M | Multilingual |
| [`small`](https://huggingface.co/vpermilp/whisper/tree/small) | [`openai/whisper-small`](https://huggingface.co/openai/whisper-small) | 244M | Multilingual |
| [`medium`](https://huggingface.co/vpermilp/whisper/tree/medium) | [`openai/whisper-medium`](https://huggingface.co/openai/whisper-medium) | 769M | Multilingual |
| [`large-v3`](https://huggingface.co/vpermilp/whisper/tree/large-v3) | [`openai/whisper-large-v3`](https://huggingface.co/openai/whisper-large-v3) | 1.55B | Multilingual |

Always select a revision. The `main` branch is an index and does not contain weights.

Each model branch also provides an experimental `model-fp8.safetensors`. Linear matrices use OCP E4M3 with one FP16 scale per output channel. Token and positional embeddings, convolutions, biases, and normalization parameters remain FP16/FP32. FP8 reduces storage but requires runtime support for the companion `*.weight_scale` tensors.

## Transformation

The conversion is deterministic and does not train or otherwise alter the model:

- Linear, convolution, embedding, and other compute weights are stored as IEEE FP16.
- Positional embeddings and LayerNorm affine parameters remain FP32, matching Whisper's mixed-precision inference behavior.
- Hugging Face Transformers parameter paths are normalized to the concise OpenAI Whisper layout, for example `model.decoder.layers.0.self_attn.q_proj.weight` becomes `decoder.blocks.0.attn.query.weight`.
- Tied token embeddings are stored once in FP16 and used for both token lookup and vocabulary projection.
- Non-floating tensors are preserved without conversion.
- Tokenizer, generation, preprocessing, and architecture configuration files are copied from each source revision.

The exact converter is included as [`convert.py`](https://huggingface.co/vpermilp/whisper/blob/main/convert.py) and can be run without PyTorch:

```bash
uv run --with numpy --with safetensors python convert.py \
  source/model.safetensors model.safetensors \
  --source openai/whisper-medium@abdf7c39ab9d0397620ccaea8974cc764cd0953e
```

Add `--with ml-dtypes` and `--compute-dtype float8_e4m3fn` to produce the scaled FP8 variant.

| Revision | Converted size | Converted SHA-256 |
|---|---:|---|
| `tiny` | 77 MB | `bbdc114a81c7775baffe0d3f045e5ab591eb3a6e1ac00a480136ff770ea783ba` |
| `small` | 486 MB | `b02cafebb1e736ea3968de06d96041a3dfc48a22820d74f5391b7daae5e36728` |
| `medium` | 1.53 GB | `e6d6ede74e0f4fd5d00042ecc56b6c4efa9b9ad6cd93e870379e9cbd0c884435` |
| `large-v3` | 3.09 GB | `c33cd318544b49c0586589c9b9bffa72a90561dac4cb59c76fc3c31314c98e95` |

| Revision | FP8 size | FP8 SHA-256 |
|---|---:|---|
| `tiny` | 60 MB | `8d3845821e2ac5f8f6d9aa556481900d53fe3895436f217d664574295cead50c` |
| `small` | 289 MB | `e9fd7e3418781a1b3468685522d90aeebe8df729ecb1e7ecbc2de4b9a9e8540c` |
| `medium` | 829 MB | `fea31295a43f7ddbdc429dd5b7b59e5b7bebef9ee8aee7f69bfd243897c59638` |
| `large-v3` | 1.63 GB | `9806d2a748c84841f516eba5c89b00c62664706ac635fda1685caa83fb040e6c` |

On AMD gfx1151, FP8 is emulated as FP16 because the architecture has no native FP8 WMMA path. It therefore saves storage but is slower: Medium took 8.90s versus 4.70s and changed a short phrase; Large V3 took 6.20s versus 2.74s while preserving the tested transcript. Use FP16 weights for speed on gfx1151.

Validation status: Large V3 preserved the tested transcript; Medium remained coherent but was not transcript-identical; Tiny produced repetitive output on the tested Russian clip; Small has only format/structure validation. All FP8 files are therefore marked experimental rather than parity-equivalent.

The mixed storage cuts checkpoint size and prevents runtimes from repeatedly converting a large FP32 vocabulary matrix during autoregressive decoding. FP16 conversion introduces the expected rounding relative to the source FP32 checkpoints.

Because parameter names are normalized, these weights are **not a drop-in `transformers` checkpoint**. Consumers must understand the OpenAI-style paths. [Svod](https://github.com/npatsakula/svod) can load this layout directly.

## Loading

With Svod, pass the branch as the Hub revision and use dimensions matching that branch:

```rust
let model = Whisper::from_hub(
    "vpermilp/whisper",
    "medium",
    ModelDimensions::for_size(WhisperSize::Medium),
)?;
```

Select the experimental FP8 file explicitly:

```rust
let model = Whisper::from_hub_with_weights(
    "vpermilp/whisper",
    "large-v3",
    "model-fp8.safetensors",
    ModelDimensions::for_size(WhisperSize::LargeV3),
)?;
```

With the inference example, use `--weights model-fp8.safetensors`.

For generic safetensors consumers:

```python
from huggingface_hub import hf_hub_download

path = hf_hub_download(
    "vpermilp/whisper",
    "model.safetensors",
    revision="medium",
)
```

## Model

Whisper is a Transformer encoder-decoder model for multilingual speech recognition and speech translation to English. It was trained by OpenAI on 680,000 hours of weakly supervised audio data. See the paper [Robust Speech Recognition via Large-Scale Weak Supervision](https://arxiv.org/abs/2212.04356), the [original repository](https://github.com/openai/whisper), and each source model card for architecture details, evaluation results, intended uses, training data, and complete limitations.

## Limitations

These converted weights retain the source models' behavior and limitations. Whisper can hallucinate text not present in the audio, generate repetitions, perform unevenly across languages and accents, and make consequential transcription errors. Evaluate the selected model on the target language, domain, acoustic conditions, and hardware before deployment. Do not use transcription or inferred attributes for high-risk decisions, and do not transcribe people without an appropriate legal basis or consent.

No independent benchmark suite is claimed for these conversions. Results reported by OpenAI apply to the source checkpoints and evaluation procedures; FP16 rounding can produce small output differences.

## License and attribution

The source checkpoints are distributed under the Apache License 2.0. This repository preserves that license and attributes OpenAI and the Whisper authors. Model-card content is adapted in part from the OpenAI/Hugging Face Whisper cards.

```bibtex
@article{radford2022robust,
  title={Robust Speech Recognition via Large-Scale Weak Supervision},
  author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
  journal={arXiv preprint arXiv:2212.04356},
  year={2022}
}
```