--- license: cc-by-nc-4.0 library_name: transformers tags: - audio - embeddings - feature-extraction - quran - arabic datasets: - FaisaI/tadabur --- # tadabur-embedding

Tadabur Embedding Use Cases

**tadabur-embedding** is an audio embedding model pretrained on the [tadabur](https://huggingface.co/datasets/FaisaI/tadabur) dataset. It maps audio into compact vector representations — at the clip level or the frame level — that can power a wide range of downstream tasks. The model is an [EAT](https://github.com/cwx-worst-one/EAT) (Efficient Audio Transformer) base encoder (12 Transformer blocks, 768-dim) pretrained from scratch on Quranic recitation audio, then trained with a multi-axis contrastive objective that produces two specialized embedding spaces on top of the shared encoder: - **Semantic space** (384-dim) — captures *what is being recited* (ayah content). Use for retrieval, search, and deduplication. - **Speaker space** (128-dim) — captures *who is reciting* (voice identity). Use for reciter similarity and clustering. ## Use cases - **Retrieval / semantic search** — find the most similar recitations to a query clip. - **Deduplication** — detect near-duplicate audio across large collections. - **Clustering** — group recordings by reciter, style, or acoustic similarity. - **Classification** — use embeddings as features for downstream classifiers. - **Frame analysis** — use the per-frame embedding sequence (before pooling) for tasks that need temporal detail: localizing acoustic events within a clip, aligning or segmenting recitations, and detecting variations over time. ## Requirements ``` transformers == 4.40 torch torchaudio ``` If you hit `AttributeError: '...Model' object has no attribute 'all_tied_weights_keys'` (or similar) on `from_pretrained`, you're on an older cached copy of this repo's code — `transformers` versions 5.x call `self.post_init()`-dependent bookkeeping during loading that earlier revisions of this wrapper didn't set up. This has been fixed; clearing your local `transformers_modules` cache for this repo and re-downloading it again. ## Quick start ```python import torch import torchaudio from transformers import AutoModel model = AutoModel.from_pretrained("FaisaI/tadabur-embedding", trust_remote_code=True).eval() # --- Load audio (16 kHz mono) --- waveform, sr = torchaudio.load("recitation.wav") waveform = waveform.mean(dim=0, keepdim=True) # mono if sr != 16000: waveform = torchaudio.functional.resample(waveform, sr, 16000) # --- Log-mel spectrogram (EAT preprocessing) --- waveform = waveform - waveform.mean() mel = torchaudio.compliance.kaldi.fbank( waveform, htk_compat=True, sample_frequency=16000, use_energy=False, window_type="hanning", num_mel_bins=128, dither=0.0, frame_shift=10, ) # (n_frames, 128) # Pad or truncate to 1024 frames (= 10.24 s) target_length = 1024 n_frames = mel.shape[0] if n_frames < target_length: mel = torch.nn.functional.pad(mel, (0, 0, 0, target_length - n_frames)) else: mel = mel[:target_length] # Normalize with tadabur dataset statistics norm_mean, norm_std = -4.381, 3.628 mel = (mel - norm_mean) / (norm_std * 2) mel = mel[None, None] # (1, 1, 1024, 128) # --- Extract embeddings --- with torch.no_grad(): semantic = model.semantic_embedding(mel) # (1, 384) L2-normalized, ayah content speaker = model.speaker_embedding(mel) # (1, 128) L2-normalized, reciter identity features = model.extract_features(mel) # (1, 513, 768) = CLS + 512 frame patches frame_embeddings = features[:, 1:] # (1, 512, 768) frame-level (~50 Hz) ``` For **retrieval / search / deduplication**, compare `semantic_embedding` vectors with a dot product (they are L2-normalized, so this is cosine similarity). For **reciter similarity**, use `speaker_embedding` the same way. The raw encoder features from `extract_features` are best for frame-level temporal tasks and as input to downstream models. For longer audio, split into 10.24 s segments and embed each one. Frame-level embeddings (`features[:, 1:]`) preserve temporal order and can be used directly for alignment and localization tasks. ## Training data Trained on [FaisaI/tadabur](https://huggingface.co/datasets/FaisaI/tadabur), a dataset of Quranic recitation audio, in two stages: EAT self-supervised pretraining of the encoder, followed by multi-axis contrastive training (semantic axis aligned to ayah text embeddings, speaker axis trained on reciter identity with guaranteed same-reciter pairs). Spectrogram normalization statistics (`norm_mean = -4.381`, `norm_std = 3.628`) were computed on this dataset — use them (not the AudioSet defaults) when preprocessing. ## Intended use & limitations - Best suited for recitation-style Arabic speech audio; performance on unrelated audio domains is not guaranteed. - Input is expected as 16 kHz mono audio converted to 128-bin log-mel spectrograms, in windows of up to 10.24 s. ## License Released under **CC BY-NC 4.0** — free for non-commercial use with attribution. ## Acknowledgements The model architecture and pretraining recipe follow [EAT: Self-Supervised Pre-Training with Efficient Audio Transformer](https://arxiv.org/abs/2401.03497) (Chen et al., 2024). The Hugging Face wrapper code is adapted from [worstchan/EAT-base_epoch30_pretrain](https://huggingface.co/worstchan/EAT-base_epoch30_pretrain). ## Citation ```bibtex @misc{tadabur-embedding, title = {tadabur-embedding: audio embeddings for Quranic recitation}, author = {Faisal}, year = {2026}, url = {https://huggingface.co/FaisaI/tadabur-embedding} } ```