metadata
license: cc-by-4.0
pipeline_tag: audio-text-to-text
MATS: An Audio Language Model under Text-only Supervision
MATS is an audio-language multimodal Large Language Model (LLM) designed to handle multiple audio tasks using solely text-only supervision. By leveraging pre-trained audio-language alignment models such as CLAP, MATS projects the shared audio-language latent space into the LLM latent space, enabling audio comprehension without relying on audio data during training.
- Paper: MATS: An Audio Language Model under Text-only Supervision
- GitHub Repository: https://github.com/wangwen-banban/MATS
Highlights
- Text-only Supervision: Endows LLMs with audio comprehension capabilities without using audio data for fine-tuning.
- Santa Mechanism: Bridges the modality gap between audio and language embeddings within CLAP (Strongly-related noisy text with audio).
- Versatility: Capable of handling various audio tasks including captioning, classification, and Question Answering.
Citation
If you find MATS useful in your research, please cite:
@article{wang2025matsaudiolanguagemodel,
title={MATS: An Audio Language Model under Text-only Supervision},
author={Wen Wang and Ruibing Hou and Hong Chang and Shiguang Shan and Xilin Chen},
journal={arXiv preprint arXiv:2502.13433},
year={2025}
}