--- language: - en license: apache-2.0 tags: - audio - audio-classification - emotion-recognition - speech-emotion-recognition - f1 - formula1 - wavlm base_model: microsoft/wavlm-base-plus pipeline_tag: audio-classification metrics: - accuracy - f1 --- # F1 Driver Tone & Emotion Detector An advanced Speech Emotion Recognition (SER) deep learning model designed to detect the emotional tone of Formula 1 driver team radio communications in real-time. The model processes raw audio input from driver-to-pit-wall communications, extracts context-rich speech representations using a pretrained **WavLM** encoder (`microsoft/wavlm-base-plus`), and passes them through a downstream **BiLSTM with Temporal Attention mechanism** to classify driver emotions. --- ## Model Overview - **Base Audio Encoder:** [microsoft/wavlm-base-plus](https://huggingface.co/microsoft/wavlm-base-plus) (768-dimensional frozen embeddings) - **Downstream Architecture:** Bidirectional LSTM (128 hidden size) + Temporal Attention Module + Linear Classifier - **Target Emotions (6 Classes):** - `0`: **Anger** - `1`: **Disgust** - `2`: **Fear** - `3`: **Happy** - `4`: **Neutral** - `5`: **Sad** --- ## Model Architecture ```text RAW DRIVER AUDIO (.wav, .mp3, .flac) | v Audio Preprocessing Pipeline - Convert Multi-channel to Mono - Resample to 16,000 Hz (16 kHz) - Peak Amplitude Normalization | v Frozen WavLM Encoder (microsoft/wavlm-base-plus) | v Frame Embeddings [B, T, 768] | v Bidirectional LSTM (hidden_size = 128 x 2 = 256) | v Temporal Attention (Aggregates sequence into fixed context vector [B, 256]) | v Linear Classifier (256 -> 6 Logits) | v Softmax Probabilities & Emotion Output ``` --- ## Model Summary | Property | Value | | :--- | :--- | | **Model Name** | F1 Driver Tone & Emotion Detector | | **Architecture Type** | WavLM + BiLSTM + Temporal Attention | | **Base Speech Encoder** | [microsoft/wavlm-base-plus](https://huggingface.co/microsoft/wavlm-base-plus) | | **Total Parameters** | ~95.2M | | **Trainable Parameters** | ~462K (Downstream Classifier Head) | | **Frozen Parameters** | ~94.7M (Frozen Base Encoder) | | **Encoder Transformer Layers** | 12 Layers | | **Encoder Hidden Dimension** | 768 | | **Encoder Attention Heads** | 12 | | **Downstream Feature Extractor** | Bidirectional LSTM (`hidden_size` = 128) | | **BiLSTM Output Dimension** | 256 (`128 x 2`) | | **Aggregation Mechanism** | Frame-level Temporal Attention | | **Classifier Head** | Linear Layer (`256 -> 6` Logits) | | **Number of Classes** | 6 (`Anger`, `Disgust`, `Fear`, `Happy`, `Neutral`, `Sad`) | | **Target Sampling Rate** | 16,000 Hz (16 kHz Mono) | | **Inference Chunking** | 2.5s Sliding Window (1.5s Hop / 1.0s Overlap) | | **Modality** | Audio | | **Output Format** | Emotion Label, Confidence %, Probability Distribution | | **Training Loss Function** | Cross-Entropy Loss | | **Optimizer** | AdamW (`lr` = 1e-3, `weight_decay` = 1e-4) | --- --- ## Input & Output Specifications ### Direct PyTorch Model I/O (`models.py` + `best_model.pt`) For users directly loading the PyTorch `BiLSTMAttentionClassifier` model and `.pt` weights: - **Model Input:** - `x`: PyTorch Tensor of WavLM frame embeddings with shape **`[Batch_Size, Sequence_Length, 768]`** - `mask`: Timestep mask tensor with shape **`[Batch_Size, Sequence_Length]`** (Optional: 1 for valid frame, 0 for padding) - **Model Output:** - `logits`: Tensor of shape **`[Batch_Size, 6]`** (Unnormalized raw class logits) - `attn_weights`: Attention weight tensor of shape **`[Batch_Size, Sequence_Length]`** (Frame-level temporal attention scores) --- ### Full Pipeline I/O (Using `modeling_f1tone.py`) For users using the complete audio-to-prediction helper script: - **Input Requirements:** - Raw Audio File (`.wav`, `.mp3`, `.flac`, `.ogg`, `.m4a`) - Automatically converted to 16 kHz Mono and peak-normalized. - **Output Return Format:** Structured Python dictionary: ```json { "audio_file": "driver_radio_01.wav", "predicted_emotion": "Anger", "confidence": 94.25, "probabilities": { "Anger": 94.25, "Neutral": 3.10, "Disgust": 1.45, "Fear": 0.70, "Happy": 0.30, "Sad": 0.20 }, "attention_weights": [0.012, 0.045, 0.180, ...] } ``` --- ## How to Run Inference using `modeling_f1tone.py` ### Directory Setup (`input/` folder) To run predictions on audio files: 1. Create an `input/` folder in the directory where `modeling_f1tone.py` is located. 2. Place your team radio `.wav` or `.mp3` files inside `input/`. ```text ToneDetectorF1/ ├── best_model.pt ├── modeling_f1tone.py ├── config.json └── input/ ├── driver_radio_1.wav └── pit_communication_2.mp3 ``` ### Running via Command Line **Option A: Process all files in `input/` folder** ```bash python modeling_f1tone.py ``` **Option B: Pass an explicit audio file path** ```bash python modeling_f1tone.py --audio_path path/to/my_audio.wav ``` ### Python Code Integration Example ```python from modeling_f1tone import FinalPredictor # Initialize predictor with model checkpoint predictor = FinalPredictor(checkpoint_path="best_model.pt") # Perform inference on an audio clip result = predictor.predict_single("input/driver_radio_1.wav") print(f"Predicted Emotion : {result['predicted_emotion']}") print(f"Confidence : {result['confidence']:.2f}%") print("Full Probabilities:", result['probabilities']) ``` --- ## Verified Model Evaluation Results Evaluation performed on holdout test set (1,117 audio evaluation samples): | Metric | Score | | :--- | :--- | | **Test Set Accuracy** | **67.86%** | | **Test Set Macro F1-Score** | **0.6794** | | **Validation Accuracy (Epoch 8)** | **69.89%** | | **Validation Macro F1-Score** | **0.7006** | | **Feature Encoder** | `microsoft/wavlm-base-plus` (768-dim, Frozen) | | **Classifier Head** | BiLSTM (128) + Temporal Attention + Linear (6) | ### Per-Class Performance Breakdown (Test Set) | Emotion Class | Precision | Recall | F1-Score | Evaluation Support | | :--- | :--- | :--- | :--- | :--- | | **Anger** | **76.71%** | **87.96%** | **0.8195** | 191 samples | | **Neutral** | **77.85%** | **75.46%** | **0.7664** | 163 samples | | **Happy** | **79.05%** | **61.26%** | **0.6903** | 191 samples | | **Disgust** | **64.65%** | **67.02%** | **0.6581** | 191 samples | | **Fear** | **52.05%** | **66.84%** | **0.5853** | 190 samples | | **Sad** | **63.33%** | **49.74%** | **0.5572** | 191 samples | --- ## Citation & License - **License:** Apache 2.0 - **Base Encoder Paper:** *WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing* (Chen et al., 2022)