--- license: other license_name: nvidia-ngc-terms-of-use license_link: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/langid_ambernet pipeline_tag: audio-classification library_name: pytorch tags: - audio-classification - language-identification - spoken-language-identification - langid - ambernet - onnx --- # AmberNet LangID Spoken language identification across **107 languages**. Give it audio, get back the language. Runs on CPU at ~72× realtime; the only dependency is `torch` or `onnxruntime`. | | | |---|---| | Input | raw mono audio, 16 kHz, float32, any length | | Output | logits over 107 languages + a 512-d language embedding | | Parameters | 28.9 M | | Speed | 10 s of audio in **138 ms** on 4 CPU threads (ONNX, RTF 0.014) | | Architecture | depthwise-separable convs + squeeze-and-excitation, x-vector stats pooling | ## Serve with ONNX Runtime Fastest option, and the one to use in production. `ambernet.onnx` takes raw audio and has dynamic batch and length axes. ```python import json import numpy as np import onnxruntime as ort options = ort.SessionOptions() options.intra_op_num_threads = 4 session = ort.InferenceSession("ambernet.onnx", options, providers=["CPUExecutionProvider"]) labels = json.load(open("config.json"))["labels"] def identify(audio: np.ndarray) -> tuple[str, float]: """audio: float32 mono 16 kHz, shape [samples].""" lengths = np.array([audio.shape[0]], dtype=np.int64) logits, _embedding = session.run(None, {"audio": audio[None], "audio_len": lengths}) probs = np.exp(logits[0] - logits[0].max()) probs /= probs.sum() top = int(probs.argmax()) return labels[top], float(probs[top]) ``` Use `providers=["CUDAExecutionProvider"]` for GPU. The graph is plain Conv / MatMul / BatchNorm — no STFT or DFT operators — so it also loads under TensorRT and other restricted runtimes. ## Serve with PyTorch ```python import soundfile as sf import torch from modeling_ambernet import AmberNet model = AmberNet.from_pretrained(".") # returns an eval-mode nn.Module audio, sr = sf.read("speech.wav", dtype="float32") assert sr == 16000 # resample first if not print(model.classify(torch.from_numpy(audio))) # [[('en', 0.9998), ('cy', 0.0001), ('hr', 0.0000), ...]] ``` `model.to("cuda")` works as usual. `modeling_ambernet.py` needs only `torch`. ## Batching Pad clips to equal length and pass the true lengths — padded frames are masked out of every convolution and both pooling steps, so a batched result is identical to running each clip alone. ```python lengths = np.array([len(clip) for clip in clips], dtype=np.int64) batch = np.zeros((len(clips), lengths.max()), dtype=np.float32) for i, clip in enumerate(clips): batch[i, : len(clip)] = clip logits, embeddings = session.run(None, {"audio": batch, "audio_len": lengths}) ``` The 512-d embedding output is a language-space representation, usable for clustering or nearest-neighbour lookup when you need more than a label. ## Files | File | Purpose | |---|---| | `ambernet.onnx` | serving graph, raw audio → logits | | `model.safetensors` + `config.json` + `modeling_ambernet.py` | PyTorch model | | `test_ambernet.py` | self-check that ONNX and PyTorch agree | ## Languages 107 languages, from the VoxLingua107 label set. Codes are those used by the model (note the legacy codes `iw` = Hebrew, `jw` = Javanese): ``` ab af am ar as az ba be bg bn bo br bs ca ceb cs cy da de el en eo es et eu fa fi fo fr gl gn gu gv ha haw hi hr ht hu hy ia id is it iw ja jw ka kk km kn ko la lb ln lo lt lv mg mi mk ml mn mr ms mt my ne nl nn no oc pa pl ps pt ro ru sa sco sd si sk sl sn so sq sr su sv sw ta te tg th tk tl tr tt uk ur uz vi war yi yo zh ``` ## Limitations - Expects 16 kHz mono. Resample first; telephone-band (8 kHz) audio is out of domain. - Trained on YouTube speech (VoxLingua107), so it inherits that domain's accents and noise profile. - Accuracy degrades on utterances under ~5 s, on code-switching, and on singing or heavily accented speech. - Closely related languages (e.g. Bosnian/Croatian/Serbian, Hindi/Urdu) are confusable. - It always returns one of the 107 languages — there is no "unknown", "silence" or "non-speech" class. Gate on a probability threshold and run voice-activity detection upstream if that matters. ## Citation ```bibtex @article{jia2022compact, title={A Compact End-to-End Model with Local and Global Context for Spoken Language Identification}, author={Jia, Fei and Koluguri, Nithin Rao and Balam, Jagadeesh and Ginsburg, Boris}, journal={arXiv preprint arXiv:2210.15781}, year={2022} } @inproceedings{valk2021voxlingua107, title={VoxLingua107: a dataset for spoken language recognition}, author={Valk, J{\"o}rgen and Alum{\"a}e, Tanel}, booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)}, year={2021}, organization={IEEE} } ```