moonshine_tiny / README.md
IMvision12's picture
Update class name in the quick start
c667742 verified
|
Raw
History Blame Contribute Delete
3.32 kB
---
pipeline_tag: automatic-speech-recognition
license: mit
base_model: UsefulSensors/moonshine-tiny
library_name: kerasformers
tags:
- keras
- kerasformers
- moonshine
- automatic-speech-recognition
- audio
- arxiv:2410.15608
- pytorch
- jax
- tf
---
## ***See [our collection](https://huggingface.co/collections/kerasformers/moonshine-6a6acc6dae5ff619b873d6c9) for all versions of Moonshine.***
# Run Moonshine with Keras 3: JAX, PyTorch, or TensorFlow
[![GitHub](https://img.shields.io/badge/GitHub-KerasFormers-black?logo=github)](https://github.com/IMvision12/KerasFormers) [![Docs](https://img.shields.io/badge/Docs-Moonshine-blue)](https://imvision12.github.io/KerasFormers/moonshine/) [![Collection](https://img.shields.io/badge/HF-Moonshine%20collection-yellow)](https://huggingface.co/collections/kerasformers/moonshine-6a6acc6dae5ff619b873d6c9)
# kerasformers/moonshine_tiny
Paper: [Moonshine: Speech Recognition for Live Transcription and Voice Commands (arXiv:2410.15608)](https://arxiv.org/abs/2410.15608) · [HF Papers](https://huggingface.co/papers/2410.15608)
Moonshine is an English ASR encoder-decoder built for **short / live** audio: the encoder sees the raw waveform length you pass in (no Whisper-style 30 s pad), so short commands stay cheap. Output is cased and punctuated.
For more details on the model, please go to the upstream [model card](https://huggingface.co/UsefulSensors/moonshine-tiny).
Pure-**Keras 3** conversion of [`UsefulSensors/moonshine-tiny`](https://huggingface.co/UsefulSensors/moonshine-tiny) for [kerasformers](https://github.com/IMvision12/KerasFormers). One implementation runs unmodified on **TensorFlow / Torch / JAX**.
This is an **ASR** checkpoint (`MoonshineConditionalGenerate`).
## ✨ Quick start
```python
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
import soundfile as sf
from kerasformers.models.moonshine import (
MoonshineProcessor,
MoonshineConditionalGenerate,
)
model = MoonshineConditionalGenerate.from_weights("kerasformers/moonshine_tiny")
processor = MoonshineProcessor.from_weights("kerasformers/moonshine_tiny")
audio, sr = sf.read("your_audio.wav", dtype="float32") # 16 kHz mono
# Cost scales with clip length: no fixed 30 s pad like Whisper.
text = model.generate(audio, processor)
print(repr(text[0]))
```
Load any Moonshine variant the same way with `from_weights("kerasformers/<variant>")`:
| Variant | Hub |
|---|---|
| `moonshine_tiny` | [`kerasformers/moonshine_tiny`](https://huggingface.co/kerasformers/moonshine_tiny) |
| `moonshine_base` | [`kerasformers/moonshine_base`](https://huggingface.co/kerasformers/moonshine_base) |
## Tips
- Set `KERAS_BACKEND` **before** importing Keras / kerasformers.
- Prefer `MoonshineProcessor.from_weights(...)` so feature extraction matches.
- English-only; pass a list of waveforms to batch.
- See [Moonshine docs](https://imvision12.github.io/KerasFormers/moonshine/) and [Loading Weights](https://imvision12.github.io/KerasFormers/loading_weights/).
- Community / upstream safetensors still work via the `hf:` prefix, e.g. `MoonshineConditionalGenerate.from_weights("hf:UsefulSensors/moonshine-tiny")`.
## Special Thanks
A huge thank you to the Useful Sensors Moonshine authors for creating and releasing these models.
License: MIT.