license: apache-2.0
base_model: openai/whisper-large-v3
library_name: basert
pipeline_tag: automatic-speech-recognition
tags:
- basert
- apple-silicon
- whisper
- speech-recognition
whisper-large-v3
BaseRT .base build of
openai/whisper-large-v3,
OpenAI's 1.55B-parameter Whisper speech-recognition model (v3: 128 mel bins, 100 languages incl. Cantonese),
for fast local transcription on Apple Silicon.
Files
| File | Precision | Size |
|---|---|---|
whisper-large-v3-F16.base |
float16 | 3.10 GB |
whisper-large-v3-Q8.base |
8-bit linears, f16 embeddings/conv/norms | 1.68 GB |
whisper-large-v3-Q4.base |
4-bit linears, f16 embeddings/conv/norms | 987 MB |
F16 and Q8 are transcription-quality equivalent (Q8 passes the same word-error parity gates against reference openai-whisper). Q4 is the smallest and remains accurate; on some smaller variants it can occasionally repeat a word in timestamped beam decoding.
Usage
curl -LsSf https://basecompute.co/install.sh | sh
basert serve --model whisper-large-v3-F16.base
POST /v1/audio/transcriptions (multipart or JSON) returns json, text,
srt, vtt, or verbose_json (with per-segment avg_logprob /
no_speech_prob / compression_ratio / temperature and the detected
language), with optional SSE streaming and POST /v1/audio/translations. Supported
request fields: language (or "auto" to detect), prompt
(initial prompt / vocabulary bias), task (transcribe/translate). Or transcribe directly
from the CLI:
basert-transcribe whisper-large-v3-F16.base audio.wav --lang auto
Multilingual: pass language (e.g. --lang de), or auto to detect, across the 100 supported languages. task=translate produces English from any source language (also via POST /v1/audio/translations).
Released under the apache-2.0 license, inherited from the base model.