Spaces:
Sleeping
Sleeping
File size: 8,465 Bytes
c8b024e 5321105 ede2251 5321105 c8b024e 5321105 c8b024e ede2251 2c07fe1 c8b024e 5321105 c8b024e 5321105 c8b024e 5321105 c8b024e 5321105 c8b024e 5321105 c8b024e 8bb2e2e ede2251 2c07fe1 ede2251 c8b024e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 | # audio-brief API β audio in β SA3 prompt out
The Space exposes a headless endpoint that runs the analysis pipeline and
returns a ready-to-use SA3 prompt plus the measured JSON. **No wallet, no
LLM** β the prompt is the deterministic structural prompt built from
measurements alone.
- **Endpoint:** `/audio_to_prompt`
- **Base URL:** `https://kalamishere-audio-brief.hf.space/`
- **Auth:** optional shared token (see *Gating* below).
## Parameters (positional)
| # | Name | Type | Default | Notes |
|---|---|---|---|---|
| 1 | `audio` | file | β | wav/mp3/flac/ogg + mp4/mov/m4a/webm/β¦ (video & aac-family are ffmpeg-extracted) |
| 2 | `token` | str | `""` | required only when the Space secret `AUDIO_BRIEF_API_TOKEN` is set |
| 3 | `fast` | bool | `false` | skip demucs stems + bass-MIDI β ~2 s instead of ~30 s; prompt omits stem/bassline detail |
| 4 | `bpm_prior` | str | `"default"` | genre slug or numeric string to seed the beat tracker |
| 5 | `llm` | bool | `false` | run the Pollinations LLM to write a natural-language prompt (analyse β brief β lens). Needs `POLLINATIONS_API_KEY` on the Space. **Falls back to the deterministic prompt if the LLM errors** β the call never fails because of the LLM |
| 6 | `lens` | str | `"loose"` | LLM constraint: `match` (tightest, keeps section timings), `loose` (feel + BPM/key, free arrangement), `free` (mood + BPM only). Ignored when `llm=false` |
| 7 | `original_prompt` | str | `""` | the caller's current gen prompt; non-empty switches both paths to anchored steering (measured tempo/key locked, user vocabulary authoritative) |
| 8 | `embedding` | bool | `false` | also run the CLAP similarity stage and return the L2-normalized vector inline as `embedding` (float list) + `embedding_dim`. Runs as a **subprocess island** (torch/checkpoint never resident next to a gen); off by default β heavy, CPU-bound |
## Two prompt modes
- **Deterministic (default, `llm=false`)** β the prompt is measurements
formatted as structured text. No wallet, no cost, works for any caller.
- **LLM-polished (`llm=true`)** β analyse β LLM writes a brief (infers
genre/mood from the numbers) β LLM compresses it into a natural-language
SA3 prompt. Costs Pollinations pollen, billed to the Space's
`POLLINATIONS_API_KEY`. If that call fails (no key / out of pollen /
timeout) the response degrades to the deterministic prompt and says so
in `prompt_source`.
## Returns
```json
{
"prompt": "Sparse ambient micro-cue in A minor Β· BPM β¦",
"prompt_source": "llm:claude-sonnet-4-6",
"brief": "This is an extremely brief, near-static cue in A minor β¦",
"match_style_prompt": "β¦",
"embedding": [0.031, -0.052, β¦], // 512 floats, L2-normalized; null unless embedding=true
"embedding_dim": 512, // null unless embedding=true
"embedding_ckpt": "default-630k", // which CLAP checkpoint made it; don't compare across differing values
"measured": {
"bpm": 128.0, "key": "A", "key_mode": "minor", "key_confidence": 0.82,
"duration_s": 30.0, "lufs_i": -9.1, "lufs_lra": 4.2, "true_peak_db": -0.3,
"sections": [...], "stems": ["bass","drums","other","vocals"],
"stem_stats": {...}, "voiceover_present": false, "bass_midi": {...}
},
"fast_mode": false,
"stages_ok": ["bpm_key","sections","loudness","stems","bass_midi"],
"errors": []
}
```
`prompt_source` is `"llm:<model>"` on the LLM path, `"deterministic"` when
`llm=false`, or `"deterministic (llm failed: β¦)"` when an LLM run fell back.
`brief` is `null` unless `llm=true`.
## Python (recommended)
```python
from gradio_client import Client, handle_file
c = Client("kalamishere/audio-brief") # or the full .hf.space URL
# deterministic (no cost)
r = c.predict(
handle_file("track.wav"),
"YOUR_TOKEN", # "" if the Space isn't gated
True, # fast=True
"default",
False, # llm=False
"loose",
api_name="/audio_to_prompt",
)
print(r["prompt"])
# LLM-polished (Space pays pollen)
r = c.predict(
handle_file("track.wav"),
"YOUR_TOKEN", False, "default",
True, # llm=True
"loose", # lens: match | loose | free
api_name="/audio_to_prompt",
)
print(r["prompt"], "\n\nvia", r["prompt_source"])
```
## Enabling the LLM path
Set a second Space secret (Settings β Variables and secrets) so the LLM
calls have a wallet to bill:
```
POLLINATIONS_API_KEY = <your Pollinations bearer token from enter.pollinations.ai>
```
Without it, `llm=true` calls still succeed β they just fall back to the
deterministic prompt (you'll see the reason in `prompt_source`).
## Gating (protect the CPU-heavy endpoint)
The public Space runs demucs on every full call, so leave it gated. Set a
Space secret in **Settings β Variables and secrets**:
```
AUDIO_BRIEF_API_TOKEN = <a long random string>
```
When set, calls must pass a matching `token` (constant-time compared) or
get `invalid or missing API token`. When unset, the endpoint is open β
fine for local runs.
## β οΈ The LLM path is slow β use `submit()`, not `predict()`
`llm=true` makes **two** Pollinations calls (brief + prompt), so a full
request runs ~40 s. `gradio_client`'s synchronous `predict()` has a short
HTTP read timeout and will raise `httpx.ReadTimeout` on it. Use the
job/poll API instead:
```python
job = c.submit(handle_file("track.wav"), "YOUR_TOKEN", True, "default",
True, "loose", api_name="/audio_to_prompt")
while not job.done():
time.sleep(2)
r = job.result() # blocks until the queue finishes, no read timeout
print(r["prompt"], "via", r["prompt_source"])
```
Deterministic calls (`llm=false`, `fast=true`) finish in ~2 s and are fine
with plain `predict()`. If you want the LLM path to fit a single `predict()`
call, it can be collapsed to one Pollinations request (skip the separate
brief) β ~20 s and half the pollen, slightly less mood nuance.
## Similarity embedding (CLAP)
Pass `embedding=true` to get a **512-float, L2-normalized CLAP vector** back in
`embedding`. Because it's L2-normalized, cosine similarity between two tracks is
just the dot product β that's the distance measurement for "how far this track
sits from a trending cluster."
- **Deterministic on full-length audio.** laion-clap (fusion off) randomly
truncates anything longer than its ~10 s window, so a naive whole-track call
returns a *different* vector every time (cos ~0.4-0.9 between repeats of the
same file). The worker instead tiles the audio into fixed 10 s windows, embeds
each, and means them β same file β same vector (repeat cos 1.0). A self-test
repeat-embed guard fails the run if that ever regresses.
- **`embedding_ckpt`** names the checkpoint that produced the vector. Vectors
from different checkpoints are NOT comparable β gate any distance computation
on matching `embedding_ckpt`.
- **Stream-safe by construction.** CLAP is loaded only in `clap_worker.py`, a
short-lived subprocess that embeds one file and exits β torch and the
checkpoint are never resident in the API process, so an embed can't OOM or
crash an in-flight SA3 generation. A failure lands as an `errors` entry and
`embedding: null`; the rest of the response still stands.
- **Poisoned-checkpoint guard.** On load the worker embeds sine / white-noise /
silence and asserts they land far apart. A broken checkpoint (everything
collapses to one vector, top-k still *looks* plausible) fails the call loudly
instead of returning a silently-wrong vector. Threshold: `CLAP_SELFTEST_MAXCOS`
(default 0.90).
- **Checkpoint.** Defaults to the auto-downloaded general 630k-audioset
checkpoint (HTSAT-tiny). Set `LAION_CLAP_MUSIC_CKPT` to the
`music_audioset_epoch_15_esc_90.14.pt` path to use the music-specialized
checkpoint (HTSAT-base) β the worker auto-matches the architecture.
- **Cost.** CPU-bound, ~10-15 s on top of the analysis. Leave it off for runs
that only need BPM/key/sections.
## Notes / limits
- **Speed:** full analysis is demucs-bound (~30 s on HF free CPU). Use
`fast=true` for interactive/high-volume use, or upgrade the Space tier.
- **Concurrency:** the queue caps concurrency at 4 (see `demo.queue`).
Heavy parallel traffic on the free tier will swap β see the low-RAM
ceiling note in the pipeline docs.
- **YouTube isn't an API input** β pass a file. (The UI's YouTube box
yt-dlp's to a file first; that flow is UI-only.)
|