Instructions to use Modotte/AIRealNet-Audio with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Modotte/AIRealNet-Audio with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Modotte/AIRealNet-Audio", trust_remote_code=True)# Load model directly from transformers import AutoModelForAudioClassification model = AutoModelForAudioClassification.from_pretrained("Modotte/AIRealNet-Audio", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModelForAudioClassification
model = AutoModelForAudioClassification.from_pretrained("Modotte/AIRealNet-Audio", trust_remote_code=True, device_map="auto")Modotte
Overview
This is the future iteration of AIRealNet
In an era of rapidly advancing AI-generated speech, voice cloning, and audio deepfakes, the need for reliable detection tools has never been higher. AIRealNet-Audio is a binary audio classifier designed to distinguish AI-generated / spoofed audio from real human speech.
It is built on a Wav2Vec-based audio encoder that processes audio in fixed 12-second chunks. The model uses a default decision threshold of 50%, which can be adjusted based on your use case.
A key design choice addresses a common failure mode of public deepfake detectors: models quickly learn a shortcut based on embedding vector length (magnitude). To prevent this, all embeddings from the base model are L1-normalized onto the unit hypersphere, forcing the classifier to rely purely on angular (directional) information rather than magnitude.
- Class 0: AI-generated / spoof audio
- Class 1: Real human audio
Default decision threshold: 0.5 (tune per use-case).
Architecture
This is a Wav2Vec-based audio encoder operating at 16 kHz, paired with a projection layer that maps its base features of dimension 768 (pooled from shape [T, 768]) down to 256. This 256-dimensional representation is the expected input shape for the classification head, which is a single-layer MLP producing 2 output logits (with softmax applied during inference).
The core innovation of this architecture is the dual L1 unit-norm applied to the model features:
- The base features are L1-normalized once when passed to the projector.
- The projected features are L1-normalized again when passed to the classification head.
This dual-hypersphere design completely eliminates size (magnitude) bias.
Why Dual L1 Normalization?
Most studies show that deepfake embeddings tend to cluster in a specific region of feature space and exhibit a characteristic vector magnitude. As a result, many detectors learn a shortcut: they simply detect the length of the embedding vector (or its presence in a particular cluster) instead of learning meaningful acoustic or spectral cues.
To eliminate this shortcut, we force every feature vector onto the unit hypersphere via L1 normalization. Consequently:
- The classification head has only one meaningful signal left to learn from โ the direction (ฮธ) of the vector.
- Because all vectors lie on the hypersphere, the opportunity for models to exploit magnitude-based or region-specific clustering is drastically reduced.
This approach is similar in spirit to GenD, which applied a single L1 normalization onto the unit hypersphere. AIRealNet-Audio extends the idea by applying the normalization twice (dual hypersphere), once before the projector and once before the head.
Training Behavior
When training without L1 normalization we observed a consistent pattern: the model performs well in the early stages, then begins to degrade after a fixed number of steps โ even after extensive hyperparameter tuning and aggressive learning-rate reduction aimed at slower convergence.
With the dual-hypersphere architecture the opposite occurs:
- Steady upward trend in performance.
- The model continues to improve with each successive step instead of collapsing after a few thousand steps.
Loss per step
Accuracy per 1000 steps
Training Data
- AI-generated speech produced by more than 100 different TTS / voice-cloning systems.
- Large collection of real human speech drawn from varied recording conditions and sources.
- On-the-fly augmentations including multiple compression codecs, bitrate changes, and additive noise.
- Separate balanced evaluation and development sets used for monitoring.
The combination of high TTS diversity and aggressive augmentation is intended to force the model to learn genuine synthesis artifacts rather than dataset-specific fingerprints.
Limitations
- Very short utterances (significantly under 12 s) must be padded or repeated; performance on extremely short clips may be lower.
- Highly adversarial or โnano-editโ modifications of real audio remain challenging.
- Completely unseen generators or extreme domain shifts (e.g., heavy telephony distortion not seen during training) can still reduce accuracy.
- The model is a probabilistic detector; it should not be used as the sole evidence in high-stakes forensic or legal settings without human review.
Performance
Across a range of standard and challenging evaluation sets the model demonstrates strong generalization. In most evaluation datasets AIRealNet-Audio maintains an Equal error rate (EER) well below 6%.
We eliminated audio files below 3sec as model is trained on 12-seconds.
| Evaluation Dataset | EER (%) |
|---|---|
| ASVspoof5 | 1.21 |
| Voxness | 1.48 |
| Malaad | 2.11 |
| ASVspoof2019 DF | 3.14 |
| In-the-Wild | 3.64 |
Additional training-time observations
- After 2 epochs: ~0.99 accuracy on both the held-out evaluation set and the development set.
- Subsequent checkpoints continue to show steady gains in accuracy (see accuracy curve above).
- Training without the dual L1 normalization exhibits the classic early-peak-then-degradation pattern; the dual-hypersphere design removes this failure mode.
Note: Extremely high numbers on controlled evaluation sets are expected. Real-world performance depends on the distribution of generators and acoustic conditions encountered at deployment time. Always calibrate the decision threshold on data that matches your target domain.
Usage
Sample Audio
The following sample was generated by Gemini-3.8 Flash TTS (AI-generated speech):
Quick Inference
from transformers import pipeline
pipe = pipeline(
"audio-classification",
model="Modotte/AIRealNet-Audio",
trust_remote_code=True
)
result = pipe(
"https://cdn-uploads.huggingface.co/production/uploads/677fcdf29b9a9863eba3f29f/C2llrmFhlWx-oryC9wF6f.wav"
)
print(result)
Expected output:
[
{'score': 0.9979, 'label': 'AIVoice'},
{'score': 0.0021, 'label': 'HumanVoice'}
]
| Label | Meaning |
|---|---|
AIVoice |
AI-generated / spoof audio |
HumanVoice |
Real human speech |
Default decision threshold is 0.5. You can adjust it according to your precision/recall needs.
Notes for Production Use
- Resample input to 16 kHz mono.
- Segment long recordings into 12-second chunks (the training chunk size).
- Aggregate chunk-level scores (mean, max, or calibrated fusion) as needed.
Intended Use
- Detection of AI-generated / spoofed speech on social media, messaging platforms, call centers, and research datasets.
- Assistance for content moderators, journalists, fact-checkers, and platform trust-and-safety teams.
- Research baseline for audio deepfake detection under a hyperspherical feature constraint.
Not intended as the sole source of evidence in legal, forensic, or high-stakes verification scenarios without corroborating human analysis.
Ethical Considerations
- Training data construction followed the same privacy-first principles used for the original AIRealNet image model (no personal or sensitive recordings).
- Users should treat model scores as one signal among many and always apply human review near the decision threshold.
- The model card explicitly documents the dual-hypersphere design and the known length-bias failure mode so that downstream users understand both the strengths and the residual risks.
How It Works
- Input audio is resampled to 16 kHz and segmented into 12-second chunks.
- A Wav2Vec encoder extracts frame-level features [T, 768].
- Features are mean-pooled, L1-normalized, and projected to 256 dimensions.
- The 256-dimensional vector is L1-normalized a second time and passed through a linear classification head.
- Softmax yields class probabilities; a threshold (default 0.5) produces the final binary decision.
Future Work
- Improve robustness to very short utterances and adversarial nano-edits.
- Expand coverage to additional languages and emerging TTS / voice-conversion systems.
- Investigate multi-modal (audio + video / metadata) detection.
- Explore variable-length or streaming-friendly variants of the dual-hypersphere constraint.
Citation
@misc{Modotte_AIRealNet_Audio_2025,
title = {AIRealNet-Audio: Dual-Hypersphere Constrained Wav2Vec for Detecting AI-Generated vs Real Speech},
author = {Parvesh Rawal},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/Modotte/AIRealNet-Audio}
}
Acknowledgments
Special thanks to Sujal for performing all the model evaluations.
References
- Yermakov, A., Cech, J., Matas, J., & Fritz, M. (2026). Deepfake Detection that Generalizes Across Benchmarks (GenD). Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV).
arXiv:2508.06248 ยท GitHub - Microsoft / Facebook Wav2Vec 2.0 and related self-supervised speech models.
- AIRealNet โ the image counterpart that motivated this audio extension.
- Downloads last month
- 54
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Modotte/AIRealNet-Audio", trust_remote_code=True)