File size: 1,449 Bytes
9147e83
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
---
license: cc-by-4.0
pipeline_tag: audio-text-to-text
---

# MATS: An Audio Language Model under Text-only Supervision

MATS is an audio-language multimodal Large Language Model (LLM) designed to handle multiple audio tasks using solely text-only supervision. By leveraging pre-trained audio-language alignment models such as CLAP, MATS projects the shared audio-language latent space into the LLM latent space, enabling audio comprehension without relying on audio data during training.

- **Paper:** [MATS: An Audio Language Model under Text-only Supervision](https://huggingface.co/papers/2502.13433)
- **GitHub Repository:** [https://github.com/wangwen-banban/MATS](https://github.com/wangwen-banban/MATS)

## Highlights

- **Text-only Supervision:** Endows LLMs with audio comprehension capabilities without using audio data for fine-tuning.
- **Santa Mechanism:** Bridges the modality gap between audio and language embeddings within CLAP (Strongly-related noisy text with audio).
- **Versatility:** Capable of handling various audio tasks including captioning, classification, and Question Answering.

## Citation

If you find MATS useful in your research, please cite:

```bibtex
@article{wang2025matsaudiolanguagemodel,
  title={MATS: An Audio Language Model under Text-only Supervision},
  author={Wen Wang and Ruibing Hou and Hong Chang and Shiguang Shan and Xilin Chen},
  journal={arXiv preprint arXiv:2502.13433},
  year={2025}
}
```