--- pipeline_tag: automatic-speech-recognition license: mit base_model: UsefulSensors/moonshine-base library_name: kerasformers tags: - keras - kerasformers - moonshine - automatic-speech-recognition - audio - arxiv:2410.15608 - pytorch - jax - tf --- ## ***See [our collection](https://huggingface.co/collections/kerasformers/moonshine-6a6acc6dae5ff619b873d6c9) for all versions of Moonshine.*** # Run Moonshine with Keras 3: JAX, PyTorch, or TensorFlow [![GitHub](https://img.shields.io/badge/GitHub-KerasFormers-black?logo=github)](https://github.com/IMvision12/KerasFormers) [![Docs](https://img.shields.io/badge/Docs-Moonshine-blue)](https://imvision12.github.io/KerasFormers/moonshine/) [![Collection](https://img.shields.io/badge/HF-Moonshine%20collection-yellow)](https://huggingface.co/collections/kerasformers/moonshine-6a6acc6dae5ff619b873d6c9) # kerasformers/moonshine_base Paper: [Moonshine: Speech Recognition for Live Transcription and Voice Commands (arXiv:2410.15608)](https://arxiv.org/abs/2410.15608) · [HF Papers](https://huggingface.co/papers/2410.15608) Moonshine is an English ASR encoder-decoder built for **short / live** audio: the encoder sees the raw waveform length you pass in (no Whisper-style 30 s pad), so short commands stay cheap. Output is cased and punctuated. For more details on the model, please go to the upstream [model card](https://huggingface.co/UsefulSensors/moonshine-base). Pure-**Keras 3** conversion of [`UsefulSensors/moonshine-base`](https://huggingface.co/UsefulSensors/moonshine-base) for [kerasformers](https://github.com/IMvision12/KerasFormers). One implementation runs unmodified on **TensorFlow / Torch / JAX**. This is an **ASR** checkpoint (`MoonshineConditionalGenerate`). ## ✨ Quick start ```python import os os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow" import soundfile as sf from kerasformers.models.moonshine import ( MoonshineProcessor, MoonshineConditionalGenerate, ) model = MoonshineConditionalGenerate.from_weights("kerasformers/moonshine_base") processor = MoonshineProcessor.from_weights("kerasformers/moonshine_base") audio, sr = sf.read("your_audio.wav", dtype="float32") # 16 kHz mono # Cost scales with clip length: no fixed 30 s pad like Whisper. text = model.generate(audio, processor) print(repr(text[0])) ``` Load any Moonshine variant the same way with `from_weights("kerasformers/")`: | Variant | Hub | |---|---| | `moonshine_tiny` | [`kerasformers/moonshine_tiny`](https://huggingface.co/kerasformers/moonshine_tiny) | | `moonshine_base` | [`kerasformers/moonshine_base`](https://huggingface.co/kerasformers/moonshine_base) | ## Tips - Set `KERAS_BACKEND` **before** importing Keras / kerasformers. - Prefer `MoonshineProcessor.from_weights(...)` so feature extraction matches. - English-only; pass a list of waveforms to batch. - See [Moonshine docs](https://imvision12.github.io/KerasFormers/moonshine/) and [Loading Weights](https://imvision12.github.io/KerasFormers/loading_weights/). - Community / upstream safetensors still work via the `hf:` prefix, e.g. `MoonshineConditionalGenerate.from_weights("hf:UsefulSensors/moonshine-base")`. ## Special Thanks A huge thank you to the Useful Sensors Moonshine authors for creating and releasing these models. License: MIT.