File size: 2,034 Bytes
d7ae26c
 
901de71
 
 
 
 
 
 
 
d7ae26c
901de71
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2c32a94
901de71
 
 
2c32a94
901de71
 
 
2c32a94
 
901de71
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
---
license: apache-2.0
language:
- en
pipeline_tag: text-to-audio
tags:
- spatial-audio
- ambisonics
- diffusion
- arxiv:2608.29549
---

# PhysWave

Pretrained checkpoints for **PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation** (EMNLP 2026).

[Code and instructions](https://github.com/lingfengyao/PhysWave) · [Paper](https://arxiv.org/abs/2608.29549) · [Demos](https://lingfengyao.github.io/PhysWave/)

PhysWave generates single-source spatial audio from an acoustic caption and waypoints, 
or a natural-language instruction parsed by an API. Output is approximately 10 seconds
of 16 kHz first-order ambisonic audio in W, X, Y, Z channel order.
Use an ambisonic decoder for spatial listening.

## Checkpoints

Both files are required:

| File | Model |
| --- | --- |
| `physwave.ckpt` | Text- and waypoint-conditioned diffusion model |
| `vae.ckpt` | Four-channel audio VAE |

## Usage

Follow the installation instructions in the code repository. From its root, download the weights and model metadata:

```bash
python -m pip install huggingface_hub
hf download 10wind/PhysWave config.json physwave.ckpt vae.ckpt --local-dir checkpoints
python -m physwave --condition examples/telephone_static.json --output outputs/telephone.wav
```

`config.json` describes the audio format and checkpoint files. Inference uses the two `.ckpt` files through the code repository.

See the [code README](https://github.com/lingfengyao/PhysWave#natural-language-input)
for Gemini/OpenRouter configuration and more examples.

## Citation

If you use PhysWave in your research, please cite:

```bibtex
@inproceedings{yao2026physwave,
  title={PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation},
  author={Yao, Lingfeng and Huang, Chenpei and Yang, Xingke and Geng, Ziye and Luo, Changqing and Wang, Hao and Liu, Jiang and Pan, Miao},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}
```