Spaces:
Running
Running
File size: 11,514 Bytes
9fd54d0 330f477 9fd54d0 330f477 9fd54d0 330f477 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 | ---
title: OpenMusic Analysis API
emoji: π΅
colorFrom: blue
colorTo: purple
sdk: docker
pinned: false
---
# OpenMusic Music Analysis API
Versioned server-side analysis for separate global audio, temporal audio and lyrics
representations. The service deliberately does not fuse these embedding spaces.
## Architecture
```text
multipart HTTP request
β
validation + temporary-file lifetime
β
MusicAnalysisService
βββ GlobalAudioAnalyzer ββββ
βββ TemporalAudioAnalyzer ββΌββ shared CLAP audio encoder
βββ LyricsAnalyzer βββββββββββ BGE-M3 text encoder
βββ ModelRegistry
β
versioned typed response
```
The HTTP layer validates, calls the application service and serializes. Analyzer
implementations own preprocessing, inference, aggregation and normalization. Model
objects are lazy singletons behind an inference semaphore; they are never created per
request or exposed to the HTTP layer.
`AnalysisContext` decodes an upload once to mono float32 PCM at 48 kHz and caches
additional deterministic in-memory resamples by sample rate. Source-file decoding is
centralized in ffmpeg and supports mp3, m4a, flac, wav, aac, ogg and opus.
## Models
| Representation | Model | Exact revision | Dimension | Normalization |
|---|---|---|---:|---|
| `audio.global` | `laion/clap-htsat-unfused` | `8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a` | 512 | L2 per window, mean, L2 final |
| `audio.temporal` | `laion/clap-htsat-unfused` | `8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a` | 512 per segment | L2 per segment |
| `lyrics.global` | `BAAI/bge-m3` | `5617a9f61b028005a4858fdac845db406aefb181` | 1024 | L2 per chunk, token-weighted mean, L2 final |
CLAP is the existing backend's baseline, now pinned instead of following Hugging Face
`main`. BGE-M3 was selected for its MIT license, English/Russian and broader
multilingual support, semantic retrieval training, 1024-dimensional dense space and
long-document support up to 8192 tokens. Lyrics still use structural chunks so section
information is not discarded and section embeddings can be exposed later. Compared
alternatives were multilingual E5 large (good multilingual quality but a 512-token
limit), multilingual MPNet (older, 128-token sentence/paragraph setup), and GTE
multilingual base (smaller, long-context, but requires repository remote code).
Every registry item declares `model_id`, exact `model_version`, modality, dimension,
dtype, normalization, license, preprocessing version and full representation config.
The preprocessing version contains a deterministic hash of all result-affecting config,
so changing window/hop/chunk settings changes cache identity.
## Audio preprocessing
The decoder runs ffprobe validation, then one deterministic ffmpeg decode (`-threads 1`)
to mono little-endian float32 PCM at 48,000 Hz. Future model sample rates are produced
from that request-scoped PCM with deterministic polyphase resampling.
`audio.global` defaults:
- 10-second analysis windows;
- adaptive count: one for short tracks, up to four as duration grows;
- window starts are uniformly distributed over 10%β90% of the valid start range;
- short tracks are repeated deterministically to a full model window;
- every CLAP vector is L2-normalized, vectors are averaged, then normalized again.
This prevents a global track vector from effectively describing only the intro.
Windows longer than CLAP's native 10 seconds are deterministically divided into evenly
placed 10-second encoder subwindows and aggregated; CLAP never receives a long input
from which its processor could choose a random crop.
`audio.temporal` defaults:
- 10-second window and 10-second hop;
- explicit tail window so the ending is represented;
- at most 24 segments, with deterministic uniform selection if the candidate count is
larger;
- chronological millisecond timestamps and one normalized CLAP vector per segment.
The temporal summary reports only representation geometry:
`number_of_segments`, cosine `mean_adjacent_distance`, `max_adjacent_distance`,
`trajectory_variance`, and `largest_transition_index`. It does not claim to measure
emotion, climax or tension.
## Lyrics preprocessing
`LyricsPreprocessor` applies Unicode NFKC, normalizes line endings, removes control
characters, LRC metadata/timestamps and pure technical URL lines, while preserving
punctuation, paragraphs and repeated choruses. It recognizes English and Russian
section markers such as Verse/Chorus/Bridge and ΠΡΠΏΠ»Π΅Ρ/ΠΡΠΈΠΏΠ΅Π²/ΠΡΠΈΠ΄ΠΆ.
Sections are the first chunk boundary. Oversized sections are split by lines and then by
token IDs, never by arbitrary character count. The default chunk limit is 512 tokens.
Repeated sections remain repeated and therefore retain their weight. BGE-M3 uses its
document-side CLS dense representation; chunk vectors are normalized and aggregated by
token count. The internal chunk/section pipeline is ready for a later `lyrics.sections`
or `lyrics.temporal` response without changing `lyrics.global`.
## Installation and deployment
The production image uses Python 3.11. A system ffmpeg/ffprobe is required.
```bash
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn main:app --host 0.0.0.0 --port 7860
```
Or:
```bash
docker compose up --build
```
Docker Compose exposes `http://localhost:8000`; Hugging Face Spaces uses container port
7860. Model files are cached in the `hf_cache` volume.
Device selection is centralized. `OPENMUSIC_DEVICE=auto` chooses CUDA, then Apple MPS,
then CPU. Set `OPENMUSIC_DEVICE=cpu` for a forced CPU fallback. Output embeddings remain
float32 regardless of device. `OPENMUSIC_INFERENCE_CONCURRENCY=1` is the safe default
that avoids duplicate loading and unbounded concurrent GPU work.
## API
### Models
```bash
curl http://localhost:7860/v1/models
```
`GET /v1/models` is the client source of truth. `loaded` distinguishes an available lazy
analyzer from one whose model is currently resident.
### Analyze a track
```bash
curl -X POST http://localhost:7860/v1/tracks/analyze \
-F 'audio=@song.flac' \
-F 'lyrics=[Verse]
Hello world' \
-F 'track_id=local:42' \
-F 'content_identity=sha256:...' \
-F 'requested_representations=audio.global' \
-F 'requested_representations=lyrics.global'
```
Repeated fields, a comma-separated value, or a JSON array are accepted for
`requested_representations`. With no field, both audio representations are calculated,
plus `lyrics.global` only when non-empty lyrics are supplied. Explicitly requesting
`lyrics.global` without lyrics is an error.
Abbreviated response:
```json
{
"schema_version": "1",
"track": {"track_id": "local:42", "content_identity": "sha256:..."},
"representations": {
"audio.global": {
"representation": "audio.global",
"model_id": "laion/clap-htsat-unfused",
"model_version": "8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a",
"preprocessing_version": "audio-clap-global-v1.<config-hash>",
"modality": "audio.global",
"dimension": 512,
"dtype": "float32",
"normalized": true,
"configuration": {},
"embedding": [0.01, -0.02],
"analysis": {"duration_ms": 213000, "windows_used": 4}
},
"audio.temporal": {
"representation": "audio.temporal",
"model_id": "laion/clap-htsat-unfused",
"model_version": "8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a",
"preprocessing_version": "audio-clap-temporal-v1.<config-hash>",
"modality": "audio.temporal",
"dimension": 512,
"dtype": "float32",
"normalized": true,
"configuration": {},
"segments": [{"start_ms": 0, "end_ms": 10000, "embedding": [0.01]}],
"summary": {"number_of_segments": 1, "mean_adjacent_distance": 0.0,
"max_adjacent_distance": 0.0, "trajectory_variance": 0.0,
"largest_transition_index": null}
}
}
}
```
Errors never contain stack traces:
```json
{"error": {"code": "AUDIO_DECODE_FAILED", "message": "The uploaded audio could not be decoded."}}
```
`GET /v1/status` reports the selected device and loaded/available representations without
environment details. `GET /health` is a lightweight container probe.
## Configuration
| Variable | Default |
|---|---:|
| `OPENMUSIC_GLOBAL_WINDOW_SECONDS` | `10` |
| `OPENMUSIC_GLOBAL_WINDOWS` | `4` |
| `OPENMUSIC_TEMPORAL_WINDOW_SECONDS` | `10` |
| `OPENMUSIC_TEMPORAL_HOP_SECONDS` | `10` |
| `OPENMUSIC_TEMPORAL_MAX_SEGMENTS` | `24` |
| `OPENMUSIC_CLAP_BATCH_SIZE` | `4` |
| `OPENMUSIC_LYRICS_CHUNK_TOKENS` | `512` |
| `OPENMUSIC_MAX_UPLOAD_BYTES` | `104857600` |
| `OPENMUSIC_MAX_LYRICS_CHARACTERS` | `100000` |
| `OPENMUSIC_MAX_AUDIO_SECONDS` | `1800` |
| `OPENMUSIC_REQUEST_TIMEOUT_SECONDS` | `300` |
Changing result-affecting representation configuration changes its generated
`preprocessing_version` and benchmark cache key.
## Benchmark tools
Place optional UTF-8 lyrics beside audio using the same stem (`song.flac` + `song.txt`).
```bash
python tools/nearest_neighbors.py ./tracks \
--representation audio.global --top-k 10 --output reports/audio-global.json
python tools/nearest_neighbors.py ./tracks \
--representation lyrics.global --output reports/lyrics-global.json
python tools/nearest_neighbors.py ./tracks \
--representation audio.temporal --output reports/audio-temporal.json
```
The tool writes human-readable JSON and CSV. Its cache identity includes audio SHA-256,
lyrics SHA-256, model ID, exact revision, preprocessing version, representation and full
configuration. Unchanged embeddings are reused. Temporal reports contain two explicitly
experimental metrics: normalized-time interpolation with mean cosine, and classic DTW
over cosine distance. Neither is treated as the final similarity definition.
Measure the current machine with a representative track:
```bash
python tools/profile_analysis.py song.flac --lyrics song.txt
```
It reports total model startup, each representation latency, peak process RSS and peak
CUDA allocation when CUDA is available.
## Tests
```bash
pip install -r requirements-dev.txt
pytest
```
Tests use deterministic fake encoders and do not download model weights. They cover API
selection/errors, global aggregation, temporal segmentation/limits, request-scoped
decode reuse, English/Russian structured lyrics and long chunking, model registry
consistency, inference failure sanitization, and both temporal similarity baselines.
## Extension points and limitations
A future MERT analyzer should implement the audio analyzer protocol, request its sample
rate from `AnalysisContext`, declare a separate registry item and use explicit names such
as `audio.mert.global` and `audio.mert.temporal`. It must not be averaged with CLAP. It is
not enabled in V1 to avoid adding a second large model/dependency path before comparative
benchmarks exist.
- CLAP temporal vectors are a trajectory through CLAP's audio space, not a trained music
emotion representation.
- BGE-M3 semantic lyrics similarity is not lyrics emotional similarity.
- Embeddings from CLAP and BGE-M3 (and future MERT models) are different spaces and must
never be compared directly with cosine or added together.
- Temporal interpolation and DTW are baselines for experiments, not validated music
structure metrics.
- Exact revisions and preprocessing config make inference reproducible, but very small
floating-point differences can still occur across PyTorch/device/hardware versions.
|