File size: 9,752 Bytes
1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 8b0cb9e e955659 1eeaf8a e955659 1eeaf8a e955659 1eeaf8a e955659 8b0cb9e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | # Data
SAID uses reader-provided datasets through three explicit interfaces. Dataset
files remain under the terms published by their respective rights holders.
The training command validates the selected interface before constructing a
model or optimizer.
## Training data at a glance
All local paths are configured once in `configs/data.yaml`. Absolute paths are
accepted; relative paths are resolved from the directory containing that file.
| Training recipe | Local data | Fields to set | Preparation or training command |
|---|---|---|---|
| DCASE fine-tuning | Extracted DCASE2026 Task 3 development audio and labels | `dcase_recordings.root` | `said train --config configs/training/dcase_passt.yaml` |
| Audio2Sph pretraining | Original VCTK v0.80 corpus and an empty output directory | `simulated_scenes.vctk.source_root`, `simulated_scenes.vctk.prepared_root` | `said prepare --config configs/training/audio2sph.yaml` |
| Complete SAID SourceBank training | Local class-labeled clips and a SourceBank manifest | `simulated_scenes.sourcebank.manifest`, `simulated_scenes.sourcebank.source_audio_root` | `said train --config configs/training/sourcebank_passt.yaml` |
The AudioMAE recipes replace `passt` with `audiomae`. Each command validates
only the local files used by its selected data adapter, so DCASE fine-tuning
does not require VCTK or SourceBank files.
## DCASE2026 recordings
Obtain the DCASE2026 Task 3 Track A development data from the official
distribution and preserve its directory names. Evaluation and DCASE
fine-tuning use these two archives from the STAIRS26 record:
- [`32ch_audio_dev.zip`](https://zenodo.org/api/records/18171005/files/32ch_audio_dev.zip/content);
- [`labels_dev.zip`](https://zenodo.org/api/records/18171005/files/labels_dev.zip/content).
The official task and dataset records are:
- DCASE2026 Task 3: <https://dcase.community/challenge2026/>
- STAIRS26, DOI
[10.5281/zenodo.18171005](https://doi.org/10.5281/zenodo.18171005),
by the Sony and Tampere University contributors named in the record;
- STARSS23, DOI
[10.5281/zenodo.7880637](https://doi.org/10.5281/zenodo.7880637),
by the Sony and Tampere University contributors named in the record.
STAIRS26 contains the two packaged development archives consumed by SAID;
STARSS23 is retained here as an upstream data citation. The cited Zenodo
records declare the MIT License for their corresponding deposited versions.
Users retain the notices supplied with the exact files they download. The
dataset is not bundled with SAID.
```text
$HOME/datasets/dcase2026_task3/
βββ eigen_dev/
β βββ dev-train-sony/*.wav
β βββ dev-train-tau/*.wav
β βββ dev-test-sony/*.wav
β βββ dev-test-tau/*.wav
βββ labels_dev/
βββ dev-train-sony/*_std.json
βββ dev-train-tau/*_std.json
βββ dev-test-sony/*_std.json
βββ dev-test-tau/*_std.json
```
For evaluation, pass `$HOME/datasets/dcase2026_task3` directly to
`said evaluate`. For training, set the same directory as
`dcase_recordings.root` in `configs/data.yaml`. The adapter verifies
paired recording and label stems, 32-channel Eigenmike input, class indices,
and official 10 Hz frame indices. It selects capsules 6, 10, 26, and 22 and
constructs the four azimuth-rotation views described in the paper.
## VCTK for Audio2Sph pretraining
Install the Online Scene Generation dependencies before preparing data or
starting simulated-scene training:
```bash
pip install -e '.[render]'
```
The paper uses VCTK release 0.80 with 109 speakers. Obtain this release from the
[official University of Edinburgh corpus record](https://datashare.ed.ac.uk/handle/10283/2651).
Its accompanying database license is ODC Attribution License 1.0. Retain the
downloaded `README` and `COPYING` files, then place the corpus at the path
selected by `simulated_scenes.vctk.source_root`:
```text
/path/to/vctk-v0.80/
βββ README
βββ COPYING
βββ wav48/
βββ p*/**.wav
```
Set `simulated_scenes.vctk.prepared_root` to an empty output directory and run:
```bash
said prepare --config configs/training/audio2sph.yaml
```
The preparation command checks the release identity in `README`, the ODC-By
1.0 notice in `COPYING`, and all 109 speaker directories. It divides the sorted
speaker list into 99 training speakers and 10 held-out speakers. Speech activity
is measured by RMS over 0.05 s windows with a 0.1 hop ratio; samples covered by
windows at or below 0.0055 RMS are removed. Each remaining waveform is repeated
or truncated to exactly two seconds at 48 kHz. The prepared directory retains
the VCTK notices and a machine-readable record of these parameters. Audio2Sph
training validates that record before Online Scene Generation begins.
## SourceBank for complete SAID training
SourceBank is a local, license-aware index over clips obtained by each reader.
Its public interface consists of a TSV or CSV manifest and an audio root.
One practical layout is:
```text
/path/to/sourcebank/
βββ audio/
β βββ female_speech_001.wav
β βββ male_speech_001.wav
β βββ ...
βββ sourcebank.csv
```
The paper recipe requires at least one authorized row for every class ID from
0 through 12. The fixed mapping is:
| ID | Class | ID | Class |
|---:|---|---:|---|
| 0 | Female speech | 7 | Door open/close |
| 1 | Male speech | 8 | Music |
| 2 | Clapping | 9 | Musical instr. |
| 3 | Telephone | 10 | Water tap/shower |
| 4 | Laughter | 11 | Bell |
| 5 | Domestic sounds | 12 | Knock |
| 6 | Walk/footsteps | | |
The manifest fields are:
| Field | Contract |
|---|---|
| `source_dataset` | Dataset or collection name |
| `source_id` | Unique stable identifier |
| `local_audio_path` | Path relative to `source_audio_root` |
| `class_id` | Zero-based DCASE class in `[0,12]` |
| `start_seconds`, `end_seconds` | Authorized segment with `0 <= start < end` |
| `license_spdx_or_uri` | SPDX identifier or stable license URI |
| `attribution` | Attribution required by the source license |
| `source_page` | Stable source or dataset page |
| `redistribution_allowed` | Explicit `true` or `false` |
| `commercial_use_allowed` | Explicit `true` or `false` |
| `training_use_allowed` | Explicit `true` or `false` |
A minimal CSV row has the following form; a `.tsv` file uses the same columns
with tab separators:
```csv
source_dataset,source_id,local_audio_path,class_id,start_seconds,end_seconds,license_spdx_or_uri,attribution,source_page,redistribution_allowed,commercial_use_allowed,training_use_allowed
local_collection,female-speech-001,female_speech_001.wav,0,0.0,2.0,LicenseRef-UserVerified,Creator or collection attribution,https://example.org/source,false,false,true
```
`source_id` values are unique. `local_audio_path` is relative to
`source_audio_root`. Users keep every time range inside its referenced clip;
the manifest validator checks numeric ordering, and selected material shorter
than two seconds is padded to the scene length. Boolean permission values are
written literally as `true` or `false`.
Permission fields use a default-deny policy. Every row must explicitly
authorize training, and `require_commercial_use: true` additionally requires
commercial-use authorization. Absolute paths and parent-directory traversal
are rejected. The manifest therefore records provenance and permission while
the audio remains in the reader's licensed local collection.
The validator enforces the declarations supplied in the manifest; it does not
determine the legal accuracy of those declarations. Users verify the source
terms and their intended use before setting the permission fields.
`require_commercial_use: false` permits a manifest to include
non-commercially licensed material. A model trained from such a manifest has
its own distribution review and does not acquire the software's MIT license.
For an independently trained commercial model, set
`require_commercial_use: true` and use a Class Feature Encoder and
initialization that also permit the intended commercial use. The official
paper checkpoints cannot be used as initialization for that route.
The paper checkpoint used an internal 122,359-row SourceBank index with SHA256
`489f2405025d64f087a76a753cb0fae1051e91df00236ae2285ca2046d664535`.
Its composition was 43,832 VCTK rows, 70,416 MUSDB18-HQ rows, 4,672 FSD50K
rows, and 3,439 verified FSDKaggle2018 rows. That index is retained as
provenance and is not distributed. It predates the public per-row permission
fields, so the public schema is a rights-aware reconstruction interface rather
than a byte-identical representation of the internal manifest.
The paper checkpoint's selected-source attribution record accompanies each
separately distributed checkpoint bundle and is summarized in
[`LICENSES/TRAINING_DATA.md`](../LICENSES/TRAINING_DATA.md).
Set these fields in `configs/data.yaml`:
```yaml
simulated_scenes:
sourcebank:
manifest: /path/to/sourcebank/sourcebank.csv
source_audio_root: /path/to/sourcebank/audio
require_commercial_use: false
```
SAID validates every manifest row before training. Online Scene Generation
then samples 1--6 labeled sources, room geometry, source regions, activity,
and additive noise according to the paper configuration.
## Data configuration
`configs/data.yaml` contains both recorded-data and simulated-scene
sections. The top-level training configuration selects one with
`data: dcase_recordings` or `data: simulated_scenes`. Paths may be absolute or
relative to the data configuration file. `audio.array: eigenmike32` selects
the Eigenmike geometry, while `capsule_indices_1based: [6, 10, 26, 22]`
records the four **1-based Eigenmike capsule numbers** used by SAID.
|