audioshield / README.md
hiten
Rename cross_dataset_evaluation to combined_dataset_evaluation; fix README references
734aaad
|
Raw
History Blame Contribute Delete
19.4 kB
metadata
title: AudioShield AI
emoji: ๐Ÿ›ก๏ธ
colorFrom: blue
colorTo: indigo
sdk: docker
app_port: 7860
pinned: false

๐Ÿ›ก๏ธ AudioShield AI

Enterprise-Grade Deepfake Audio Detection & Forensic Analysis Platform

Detect synthetic voices with 99.11% accuracy using self-supervised speech representations, temporal sequence modeling, and attention-based forensic reasoning.

Python PyTorch Transformers FastAPI React Docker HF Spaces Accuracy AUC License: MIT

๐Ÿš€ Live Demo ยท ๐Ÿ“Š Performance ยท ๐Ÿ—๏ธ Architecture


AudioShield AI landing page

๐Ÿ’ก Why AudioShield AI?

Voice cloning technology has crossed a threshold. Tools that once required studios and hours of training data can now clone a voice from seconds of audio โ€” and the output is good enough to fool the human ear. This isn't a hypothetical risk anymore; it's actively being used in financial fraud, impersonation scams, and disinformation.

Most academic deepfake detectors stop at a Jupyter notebook and an accuracy score. AudioShield AI doesn't. It's built as a complete, deployable forensic product:

  • ๐ŸŽฏ A production-trained model evaluated with rigorous, leak-free methodology
  • ๐Ÿ–ฅ๏ธ A full-stack web application with a real inference backend, not a mock UI
  • ๐Ÿ“Š An investigation-grade dashboard that explains why a clip was flagged, not just what it was flagged as
  • ๐Ÿ“„ Exportable forensic reports suitable for documentation and review

The goal was to build something closer to a real security tool than a portfolio script โ€” from raw waveform to actionable forensic verdict.


โœจ Key Features

๐Ÿ” Real-time Deepfake Detection Classifies uploaded audio as Real or Fake in under a second
๐Ÿง  WavLM Transformer Backbone Self-supervised speech foundation model captures prosody, speaker identity, and synthetic artifacts
๐Ÿ” BiLSTM Temporal Modeling Captures sequential dependencies across the embedding sequence in both directions
๐ŸŽฏ Attention Pooling Learns to weight the most forensically relevant frames instead of naive averaging
๐Ÿ“ˆ Confidence & Risk Scoring Authenticity score, risk assessment (Low/Medium/High/Critical), and calibrated probability
๐Ÿ—บ๏ธ Chronological Suspicion Heatmap Segment-by-segment breakdown showing exactly where a clip deviates from organic speech patterns
๐ŸŽง Interactive Waveform Diagnostics Playback with suspicious regions highlighted directly on the oscilloscope view
๐Ÿ•น๏ธ Forensic Auditory Playground Blind-listening challenge โ€” test your ear against the model on real/fake samples
๐Ÿ“‘ Exportable PDF Reports One-click forensic investigation reports for documentation and review
๐Ÿ—‚๏ธ Batch Verification Logs Audit trail of historical classifications with filterable risk levels
โšก FastAPI Inference Backend Model loads once at startup; serves predictions via a real API, not a notebook loop
๐Ÿณ Dockerized Deployment Single-container deployment to Hugging Face Spaces, reproducible anywhere

๐Ÿ—๏ธ System Architecture

End-to-End Request Flow

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚    User     โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚  React Frontend   โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚  FastAPI Backend โ”‚
โ”‚  (Browser)  โ”‚     โ”‚  (Vite + Tailwind)โ”‚     โ”‚    (Uvicorn)     โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                                        โ”‚
                                                        โ–ผ
                                          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                                          โ”‚   Audio Preprocessing    โ”‚
                                          โ”‚  16kHz ยท Mono ยท Normalizeโ”‚
                                          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                                        โ–ผ
                                          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                                          โ”‚      deployment_model.pt  โ”‚
                                          โ”‚   (loaded once at startup)โ”‚
                                          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                                        โ–ผ
                                            Prediction + Confidence
                                                        โ”‚
                                                        โ–ผ
                                          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                                          โ”‚   JSON Response โ†’ React   โ”‚
                                          โ”‚   Forensic Dashboard      โ”‚
                                          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Model Pipeline Architecture

   Audio Input          WavLM Base+           BiLSTM            Attention Pooling        Classifier
  (WAV/MP3/FLAC)   โ”€โ”€โ–ถ  SSL Representation โ”€โ”€โ–ถ  Temporal  โ”€โ”€โ–ถ     Feature Weights   โ”€โ”€โ–ถ   Softmax
                        (Contextual Embeddings)  Context                                  (Real / Fake)
Stage Role
Preprocessing Resamples to 16kHz, converts to mono, normalizes amplitude, handles variable-length audio via padding/truncation
WavLM Base+ Microsoft's self-supervised speech foundation model โ€” extracts rich contextual embeddings capturing speaker traits, prosody, and synthetic artifacts
BiLSTM Models temporal dependencies across the embedding sequence in both forward and backward directions
Attention Pooling Learns to weight the most discriminative frames rather than pooling naively, improving sensitivity to localized synthetic artifacts
Classifier Fully connected layer with dropout + GELU activation, outputting a calibrated Real/Fake probability

Pre-trained self-supervised speech representations provide strong zero-shot generalization to unseen deepfake generation methods โ€” a key reason this architecture outperforms handcrafted-feature baselines (MFCC, Chroma, Spectrogram).


๐Ÿงฐ Tech Stack

Machine Learning

  • Python
  • PyTorch
  • Hugging Face Transformers
  • WavLM Base+
  • NumPy ยท Pandas

Audio Processing

  • Librosa
  • SoundFile

Frontend

  • React
  • Vite
  • Tailwind CSS

Backend & Deployment

  • FastAPI
  • Uvicorn
  • Docker
  • Hugging Face Spaces

๐Ÿ“Š Dataset

Trained and evaluated on a combined dataset merging Fake-or-Real (FoR) and ASVspoof 2019 LA โ€” 162,374 total samples.

Dataset Samples
ASVspoof 2019 LA 121,461
Fake-or-Real (FoR) 40,913
Combined Total 162,374
Split Samples
Training 55,584
Validation 30,919
Testing 75,871

The test set (75,871 samples) combines 71,237 ASVspoof + 4,634 FoR samples, providing a rigorous held-out evaluation across diverse attack types and recording conditions.


๐Ÿ”ฌ Audio Preprocessing

Every recording is standardized before feature extraction:

  1. Loading โ€” via Librosa / SoundFile
  2. Resampling โ€” converted to 16kHz (WavLM's expected input rate)
  3. Mono conversion โ€” removes channel-based variance
  4. Normalization โ€” amplitude-normalized so volume doesn't bias predictions
  5. Padding / Truncation โ€” handles variable-length clips during batching (max 96,000 samples / 6 sec)

๐Ÿง  Feature Extraction โ€” Why WavLM?

Traditional handcrafted features (MFCC, Chroma, Spectrograms) were deliberately avoided in favor of WavLM Base+, Microsoft's self-supervised speech foundation model.

Raw Audio โ†’ WavLM Processor โ†’ Transformer Encoder โ†’ Contextual Embeddings โ†’ BiLSTM โ†’ Attention Pooling โ†’ Fixed Vector

Why this matters: WavLM is pretrained on massive unlabeled speech corpora, learning representations of speaker characteristics, prosody, and speech dynamics that generalize far better than hand-engineered spectral features โ€” particularly against unseen deepfake generation methods, which is the failure mode that sinks most academic detectors in the real world.


๐Ÿ‹๏ธ Training Pipeline

  • Loss: Focal Loss (ฮฑ=0.25, ฮณ=2.0)
  • Optimizer: AdamW (lr=1ร—10โปโต, weight_decay=0.01)
  • Scheduler: Cosine with linear warmup (10% warmup steps)
  • Batch Size: 32
  • Max Audio Length: 6 seconds (96,000 samples at 16kHz)
  • Epochs: 7 (early stopping based on AUC improvement)
  • Mixed Precision: FP16 via GradScaler
  • Freeze Strategy: WavLM frozen for first 2 epochs, then unfrozen for fine-tuning
  • Output: deployment_model.pt โ€” the exact artifact served in production

The combined dataset was assembled by merging FoR and ASVspoof 2019 LA, then splitting into train/validation/test sets. Training was performed on Kaggle with GPU acceleration.


๐Ÿ“ˆ Diagnostic Performance Metrics

Evaluated on the combined held-out test set (75,871 samples: 71,237 ASVspoof + 4,634 FoR).

Metric Overall Real (Bonafide) Spoof (Fake)
Accuracy 99.11% โ€” โ€”
Precision โ€” 94.14% 99.87%
Recall โ€” 99.12% 99.10%
F1 Score โ€” 96.57% 99.49%
ROC-AUC 0.9994 โ€” โ€”
Equal Error Rate (EER) 0.89% โ€” โ€”

โœ… Why These Results Are Trustworthy

High numbers invite skepticism โ€” rightfully so. Here's why this isn't overfitting:

  • No data leakage โ€” splits were created before training; test data was never seen during training or model selection.
  • Large, diverse test set โ€” 75,871 held-out samples spanning two independent datasets with different attack types ensures robust evaluation.
  • Large training set โ€” 55,584 labeled samples reduce memorization risk and support generalization.
  • Transfer learning, not memorization โ€” performance is driven primarily by WavLM's pretrained representations, not example-level memorization.
  • Consistent validation/test alignment โ€” closely matched performance across splits indicates genuine generalization rather than overfitting to validation.

๐Ÿ“Š Combined Dataset Evaluation

The model was trained on a combined dataset merging Fake-or-Real (FoR) and ASVspoof 2019 LA, then evaluated on a held-out test set from the same combined distribution.

Evaluation Methodology

Component Detail
Training Datasets FoR + ASVspoof 2019 LA (combined)
Total Samples 162,374 (55,584 train / 30,919 val / 75,871 test)
Model Checkpoint deployment_model.pt โ€” best epoch by validation AUC
Preprocessing 16kHz resample, mono, head truncation to 6s, WavLM feature extraction
Evaluation Type In-distribution held-out test

Per-Class Metrics

Class Precision Recall F1-Score Support
Real (Bonafide) 94.14% 99.12% 96.57% 9,619
Spoof (Fake) 99.87% 99.10% 99.49% 66,252
Overall Accuracy 99.11% 75,871

The model achieves near-perfect spoof detection (99.87% precision, 99.10% recall) with a slight trade-off in real-class precision (94.14%), meaning it occasionally flags real audio as suspicious โ€” a conservative bias that favors security over false negatives.

How to Reproduce

cd combined_dataset_evaluation
python run_inference_for.py          # Run FoR inference
python compute_metrics_for.py        # Compute metrics + visualizations

See the model evaluation directory for full metrics, confusion matrix, ROC/PR curves, and score distributions.


๐Ÿ–ฅ๏ธ Web Application

Forensic Sandbox

Drag-and-drop upload with reference benchmark cases pulled from the FoR and ASVspoof 2019 LA datasets, plus live duration/sample-rate/inference-time readouts per file.

Forensic Sandbox upload interface

Investigation Dashboard โ€” Analysis Overview

Authenticity score, detection profile, model backbone, investigation findings, confidence distribution curve, and an acoustic anomaly radar chart (Spectral Inconsistency, Vocoder Footprint, Phase Coherence, Breath Mark Gaps, Jitter/Shimmer Ratio).

Forensic Investigation Dashboard - Analysis Overview

Forensic Metrics โ€” Chronological Suspicion Heatmap

Segment-by-segment breakdown of suspicion scores across the clip timeline, paired with an interactive waveform oscilloscope that highlights suspicious regions directly on playback.

Chronological suspicion heatmap and waveform diagnostics

Investigation Report

One-click downloadable forensic PDF report โ€” audio summary, key findings, and full timeline analysis, ready for documentation or review.

Forensic Investigation Report

Model Pipeline & Diagnostic Metrics

Model Pipeline Architecture panel Diagnostic Performance Metrics panel

๐Ÿ“ Repository Structure

Audio-Shield-AI/
โ”‚
โ”œโ”€โ”€ combined_dataset_evaluation/   # In-distribution evaluation on combined dataset
โ”‚   โ”œโ”€โ”€ download_asvspoof.py       # ASVspoof 2019 LA download script
โ”‚   โ”œโ”€โ”€ run_inference.py           # Pure inference on ASVspoof
โ”‚   โ”œโ”€โ”€ compute_metrics.py         # Metrics + visualizations
โ”‚   โ”œโ”€โ”€ error_analysis.py          # FP/FN + attack-type analysis
โ”‚   โ”œโ”€โ”€ generalization_report.py   # Honest generalization report
โ”‚   โ”œโ”€โ”€ run_all.py                 # Orchestrator
โ”‚   โ”œโ”€โ”€ asvspoof_data/             # Dataset (downloaded)
โ”‚   โ”œโ”€โ”€ results/                   # Metrics + CSVs
โ”‚   โ””โ”€โ”€ figures/                   # ROC, PR, confusion matrix, etc.
โ”‚
โ”œโ”€โ”€ notebooks/
โ”‚   โ”œโ”€โ”€ 1_eda-data-validation.ipynb
โ”‚   โ”œโ”€โ”€ 2_preprocessing.ipynb
โ”‚   โ”œโ”€โ”€ 3_Model.ipynb
โ”‚   โ”œโ”€โ”€ 4_Evaluation.ipynb
โ”‚   โ””โ”€โ”€ 5_inference.ipynb
โ”‚
โ”œโ”€โ”€ audioshield-app/              # React frontend source (Vite)
โ”‚   โ””โ”€โ”€ src/
โ”‚       โ”œโ”€โ”€ App.tsx
โ”‚       โ””โ”€โ”€ index.css
โ”‚
โ”œโ”€โ”€ static/                       # Built frontend (served by FastAPI)
โ”œโ”€โ”€ assets/                       # Screenshots and images
โ”œโ”€โ”€ deployment_model.pt           # Trained model checkpoint
โ”œโ”€โ”€ app.py                        # FastAPI server
โ”œโ”€โ”€ Dockerfile
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ .gitignore
โ””โ”€โ”€ README.md

๐Ÿš€ Getting Started

Reproduce Training (Kaggle)

The notebooks were developed entirely on Kaggle to take advantage of free GPU access and the dataset's native Kaggle hosting โ€” avoiding a ~70,000-file local download and giving anyone reproducing this work a zero-setup environment.

  1. Open notebooks/1_eda-data-validation.ipynb in a new Kaggle Notebook
  2. Attach the Fake-or-Real Audio Dataset
  3. Enable GPU acceleration (Settings โ†’ Accelerator โ†’ GPU)
  4. Run notebooks sequentially: 1 โ†’ 2 โ†’ 3 โ†’ 4

Local Deployment

Clone the repository

git clone https://github.com/gautamhardik/Audio-Shield-AI.git
cd Audio-Shield-AI

Backend

pip install -r requirements.txt
uvicorn app:app --reload --port 7860

Frontend (development)

cd audioshield-app
npm install
npm run dev

Docker (full stack)

docker build -t audioshield-ai .
docker run -p 7860:7860 audioshield-ai

Hugging Face Spaces Deployment

The production app runs as a Docker Space on Hugging Face, bundling the React frontend, FastAPI backend, and deployment_model.pt into a single reproducible container with a public URL:

๐Ÿ”— hardik-25-audioshield.hf.space


๐Ÿ”ฎ Future Work

  • ๐ŸŒ Multilingual deepfake detection
  • ๐ŸŽ™๏ธ Streaming / real-time inference for live calls
  • ๐Ÿ”Ž Explainable AI โ€” frame-level saliency visualization
  • ๐Ÿ›ก๏ธ Adversarial robustness testing against evasion attacks
  • ๐Ÿ“ฑ Mobile deployment (on-device inference)
  • ๐Ÿ“ฆ Batch analysis API for bulk investigations
  • ๐Ÿ—ฃ๏ธ Speaker attribution / voice fingerprinting

๐Ÿ”ฌ Model Evaluation

All evaluation code, results, metrics, and visualizations are available in the combined_dataset_evaluation/ and model evaluation directories.


๐Ÿ™ Acknowledgements

  • Microsoft โ€” for WavLM and the broader speech-SSL research that made this possible
  • Hugging Face โ€” Transformers library and Spaces hosting
  • PyTorch team
  • Kaggle โ€” free GPU compute and dataset hosting
  • Creators of the Fake-or-Real and ASVspoof datasets

๐Ÿ“„ License

Distributed under the MIT License. See LICENSE for details.


๐Ÿ›ก๏ธ AudioShield AI

Enterprise Deepfake Audio Detection Platform

Live Demo

If this project was useful or interesting, consider โญ starring the repo.