| --- |
| license: other |
| license_name: unresolved-upstream |
| tags: [source-separation, senselab, mirror] |
| --- |
| |
| # unasdiff diffusion source priors (mirror) |
|
|
| Mirror of the two released priors for **Unsupervised Single-Channel Audio |
| Separation with Diffusion Source Priors** (Shi et al., AAAI 2026), |
| [arXiv:2512.07226](https://arxiv.org/abs/2512.07226). |
|
|
| Held so that senselab's `source_separation` unasdiff backend can pin a revision |
| and a content hash. Nothing here is original work of the Senseable Intelligence |
| Group. |
|
|
| ## Contents |
|
|
| | File | What it is | |
| |---|---| |
| | `speech_source.pt` | Speech prior, trained on VCTK. `num_class = 1`; its only label is `0`. | |
| | `sound_source.pt` | General-sound prior, trained on FSDKaggle2018. `num_class = 50`, of which indices 0–40 are the 41 trained classes; the remaining nine slots are unused headroom. | |
| | `atten_unet_vctk.toml` | Upstream config for the speech prior. | |
| | `atten_unet_fsd.toml` | Upstream config for the sound prior. | |
|
|
| The two label spaces are **separate**. A label index is only meaningful against |
| the prior it belongs to. |
|
|
| Each checkpoint carries `model` and `ema` state dicts. Upstream's own loader |
| returns the **`ema`** copy; loading `model` instead runs but separates measurably |
| worse, without failing. |
|
|
| ## Provenance |
|
|
| - Upstream code: <https://github.com/RunwuShi/unasdiff>, pinned at |
| `5a5d70cdc94fe9d034892a1c5bc68ad1a67d2daa`. |
| - Weights: the authors' own Google Drive releases linked from the upstream README. |
|
|
| ## Licence status — unresolved |
|
|
| The upstream repository has **no `LICENSE` file and no licence statement**, so the |
| default is all rights reserved. A request for an explicit licence has been opened |
| upstream; **this mirror stays private until it is answered.** |
|
|
| ## Known limitation, from the authors |
|
|
| > The source-model-based separation approach is not well suited for same-class |
| > source separation (e.g. speech separation), because it lacks speaker-conditioning. |
|
|
| senselab exposes speech–speech separation anyway, with that caveat in the |
| docstring, because the alternative is a user rediscovering it by measurement. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @inproceedings{shi2026unasdiff, |
| title = {Unsupervised Single-Channel Audio Separation with Diffusion Source Priors}, |
| booktitle = {Proc. AAAI}, |
| year = {2026} |
| } |
| ``` |
|
|