--- license: cc-by-nc-sa-4.0 tags: - music - audio - contrastive - clap - music-information-retrieval library_name: amclap --- # AllMusicCaps: Baseline + Structured captions Baseline plus LLM-structured captions. Audio-text contrastive (CLAP) model over the OMAR-RQ audio encoder and an `all-mpnet-base-v2` text encoder, trained on music with album-review derived captions. Embeddings are 512-dimensional; audio is expected at 24 kHz. ## Usage ```python from amclap import get_model model = get_model(model_id="mtg-upf/allmusiccaps_baseline_struct", device="cpu") audio_emb = model.forward_audio(audio) # (B, 512), audio at 24 kHz text_emb = model.forward_text(["warm analog synth pads"]) # (B, 512) ``` ## Training | | | |---|---| | Text encoder | frozen (stock all-mpnet-base-v2) | | Reported checkpoint | step 149,796 | | Audio encoder layers | see `config.gin` | The text encoder was frozen during training, so stock `all-mpnet-base-v2` weights are reloaded at construction. ## License CC BY-NC-SA 4.0. See `LICENSE`. ## Citation ```bibtex @inproceedings{alonso2026allmusiccaps, title = {{AllMusicCaps}: Album Reviews as Complementary Supervision for Music {CLAP}}, author = {Alonso-Jim{\'e}nez, Pablo and Lizarraga-Seijas, Xavier and Serra, Xavier and Bogdanov, Dmitry}, booktitle = {International Society for Music Information Retrieval Conference (ISMIR)}, year = {2026}, } ```