Instructions to use blazeofchi/Aural-One-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blazeofchi/Aural-One-E2B with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("blazeofchi/Aural-One-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Aural One E2B Preview
Native audio in, structured decisions out. Aural One adapts Gemma 4 E2B to score supplied choices from a recording, a written state, and named questions. One model handles the audio and the decision; its inference path does not require speech-to-text or a separate sound classifier.
This v0.1.0 research preview publishes a frozen adapter and 54.3 million updated acoustic/projection weights. The pinned Gemma 4 E2B base is downloaded separately. The language model was frozen during this final training stage.
What we measured
| Evaluation | Aural One preview |
|---|---|
| English CREMA-D actor-held-out development, macro-F1 | 0.391 on 445 clips |
| Bangla SUBESCO speaker-held-out development, macro-F1 | 0.295 on 700 clips, up from 0.134 reference |
| Bangla SUBESCO four-speaker reserved test, macro-F1 | 0.344 on 1,012 consensus clips |
| Structured choice / No-Null / ordinal development | 189 / 174 / 191 correct out of 200 each |
| Warm pod-local HTTP, ~58-second Opus with three questions | 0.508 s p50 / 0.579 s p95 |
These numbers have different test scopes. The full evaluation report covers the frozen reference, listener-vote cross-entropy and calibration, same-words contrasts, sound events, Italian/German transfer, long audio, warm client latency, concurrency, and GPU-only cost. Release validation records the public-Hub load test. Aural One is an early release with room to improve rare emotions and natural sound transfer. The sub-second number above is measured inside the warm GPU pod; Tokyo end-to-end sub-second latency is still a research goal.
Use the model
The public GitHub repository contains the pinned loader, JSON example, full acoustic fine-tuning code and data schema, and evaluation entry point. Use Python 3.12 with an NVIDIA GPU; install a compatible PyTorch build first, then:
pip install git+https://github.com/Parassharmaa/aural-one.git
from aural_one import load_aural_one
model = load_aural_one("blazeofchi/Aural-One-E2B")
result = model.score(
audio="/absolute/path/to/your-audio.wav",
state={"task": "listen to the voice"},
questions={
"emotion": {
"question": "Which emotion is most evident in the speaker's voice?",
"options": ["angry", "disgusted", "fearful", "happy", "neutral", "sad", "surprised"],
},
"baby_cry": {
"question": "Is a baby crying audible?",
"options": ["No", "Yes"],
},
},
)
print(result)
The preview scores 2–8 options per named question. A binary or ordinal value is represented by its supplied options. Probabilities are normalized over those options and are not calibrated confidence estimates. The simple public loader scores questions separately; the warm HTTP timing above used an experimental shared-audio serving path. The reference loader passed short and synthetic 58-second smoke tests on a 24 GB Blackwell GPU partition, with 9.8 GiB PyTorch allocation. A lower minimum has not been established.
Model and data
The base revision is 3e22461f65e89153144f8adb70e3b8c2cc9845a7. The acoustic delta SHA-256 is ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96 and the adapter SHA-256 is e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3. release.json pins them for the loader.
The selected run used crowd-voted CREMA-D English acted speech, listener-voted SUBESCO Bangla acted speech, HeySQuAD Human, MInDS-14, and original synthetic speech for typed decisions. Training kept the language model and prior adapter frozen while updating the last two audio Conformer blocks and the two audio projections. Details and source terms are in TRAINING.md. No training or evaluation audio is uploaded here.
Intended use
Use this preview for research and prototyping of audio-grounded, state-conditioned choices. Acted-speech scores may not transfer uniformly to spontaneous speech, accents, or recording conditions. Avoid using its emotion judgments as high-stakes assessments of people.
Code and Aural One weight deltas are Apache 2.0. The separately downloaded Gemma 4 E2B base is also Apache 2.0.
The public release checklist records the final model-card, code, chart, hash, and GPU smoke checks.
Audio in the introduction: CC0 recording credits.
