Firebot Voice Intent Classifier & Speaker Verification Models
Official model checkpoints and speaker verification voiceprints for Firebot, an autonomous tactical firefighting robot platform with dual-tier offline voice intent recognition and personalized acoustic adaptation.
Overview
The Firebot voice system is engineered for zero-latency, high-reliability command recognition in noisy operational environments. It combines:
- Frozen Whisper Encoder Feature Extraction: 384-dimensional pooled audio embeddings from OpenAI's Whisper model (
tiny.en/base.en). - L2-SP Regularized MLP Intent Classifier: A lightweight Multi-Layer Perceptron (LayerNorm $\to$ Linear 384$\times$128 $\to$ GELU $\to$ Dropout $\to$ Linear 128$\times$15) fine-tuned with an L2 penalty pulling weights toward the base anchor.
- ECAPA-TDNN Multi-Clip Acoustic Voiceprints: 192-dimensional speaker embeddings for tactical operator authentication and biometric voice gating.
- Adaptive Online Calibration: Few-shot per-operator head adaptation enabling rapid personalization from 1β5 takes per command with held-out generalization validation.
Model Architecture & Directory Layout
firebot-voice-intent/
βββ intent_head.pt # Base 15-class Whisper intent classification head
βββ intent_head.json # Architecture configuration, pooling, and class map
βββ intent_prototypes.pt # Latent prototype embeddings for few-shot similarity
βββ users/
β βββ ananya.pt # Calibrated personal Whisper head for Operator Ananya
β βββ ananya.json # Empirical training metrics and per-class take breakdown
β βββ ananya.history.jsonl # Training iteration telemetry log
β βββ avinandan.pt # Calibrated personal Whisper head for Operator Avinandan
β βββ avinandan.json # Calibration metrics
β βββ avinandan.history.jsonl
βββ voiceprints/
βββ ananya.npy # ECAPA-TDNN biometric speaker voiceprint embedding
βββ avinandan.npy # ECAPA-TDNN biometric speaker voiceprint embedding
Command Vocabulary (15 Closed-Set Classes)
| Class ID | Canonical Label | Voice Phrases / Synonyms | Target Coordinate / Action |
|---|---|---|---|
0 |
STOP |
"stop", "halt", "freeze", "abort mission", "e-stop" | Emergency brake & valve cutoff |
1 |
EXTINGUISH |
"put out the fire", "extinguish", "spray the fire", "suppress" | Engage high-pressure water pump |
2 |
RETURN_HOME |
"return home", "go to base", "back to dock", "retreat" | Autonomous navigation to dock (1.2, 1.0) |
3 |
STATUS |
"status report", "give me a report", "tank level", "sitrep" | Telemetry & battery check |
4 |
UNKNOWN |
"hello", "what's the weather", "testing" | Out-of-domain conversational filter |
5 |
GOTO_HOME |
"go to home", "head to charging station" | Waypoint navigation: (1.2, 1.0) |
6 |
GOTO_CENTER |
"go to center", "move to the middle" | Waypoint navigation: (6.0, 4.0) |
7 |
GOTO_NORTH |
"go to north", "move to the north side" | Waypoint navigation: (6.0, 7.0) |
8 |
GOTO_SOUTH |
"go to south", "head south" | Waypoint navigation: (6.0, 1.0) |
9 |
GOTO_EAST |
"go to east", "move right" | Waypoint navigation: (10.5, 4.0) |
10 |
GOTO_WEST |
"go to west", "move left" | Waypoint navigation: (2.0, 4.0) |
11 |
GOTO_NORTHEAST |
"go to northeast", "upper right corner" | Waypoint navigation: (10.5, 7.0) |
12 |
GOTO_NORTHWEST |
"go to northwest", "upper left corner" | Waypoint navigation: (1.5, 7.0) |
13 |
GOTO_SOUTHEAST |
"go to southeast", "bottom right corner" | Waypoint navigation: (10.5, 1.0) |
14 |
GOTO_SOUTHWEST |
"go to southwest", "bottom left corner" | Waypoint navigation: (1.5, 1.2) |
Quickstart: Python Inference
import soundfile as sf
import torch
from huggingface_hub import hf_hub_download
from firebot.voice_intent.infer import IntentClassifier
# 1. Download model from Hugging Face Hub
ckpt_path = hf_hub_download(repo_id="anabaena/firebot-voice-intent", filename="intent_head.pt")
# 2. Instantiate Intent Classifier (loads frozen Whisper + lightweight head)
classifier = IntentClassifier(checkpoint_path=ckpt_path)
# 3. Classify raw audio (16kHz mono WAV)
audio, sr = sf.read("command_sample.wav", dtype="float32")
prediction = classifier.predict(audio, sample_rate=sr)
print(f"Predicted Intent: {prediction['label']}")
print(f"Confidence: {prediction['confidence']:.2%}")
print(f"Action: {prediction['canonical_phrase']}")
Biometric Speaker Verification (ECAPA-TDNN)
import numpy as np
# Load enrolled speaker voiceprint
enrolled = np.load("voiceprints/ananya.npy")
# Compare with live speaker embedding
cosine_sim = np.dot(enrolled, live_embedding) / (np.linalg.norm(enrolled) * np.linalg.norm(live_embedding))
is_authorized = cosine_sim >= 0.72 # calibrated acceptance threshold
print(f"Speaker Verified: {is_authorized} (Similarity: {cosine_sim:.3f})")
Performance & Evaluation
- Inference Latency: ~35ms on CPU (Apple Silicon / modern x86_64)
- Base Intent Accuracy: 94.2% on diverse acoustic validation test set
- Personalized Operator Accuracy: 98.8% β 100.0% with 3β5 takes per command
- Memory Footprint: ~210 KB for the MLP head, ~75 MB for frozen Whisper tiny encoder
License
MIT License. Designed and developed as part of the Firebot Autonomous Robotic Platform.
Evaluation results
- Accuracyself-reported0.962