Spaces:
Sleeping
Sleeping
| # audio-brief API β audio in β SA3 prompt out | |
| The Space exposes a headless endpoint that runs the analysis pipeline and | |
| returns a ready-to-use SA3 prompt plus the measured JSON. **No wallet, no | |
| LLM** β the prompt is the deterministic structural prompt built from | |
| measurements alone. | |
| - **Endpoint:** `/audio_to_prompt` | |
| - **Base URL:** `https://kalamishere-audio-brief.hf.space/` | |
| - **Auth:** optional shared token (see *Gating* below). | |
| ## Parameters (positional) | |
| | # | Name | Type | Default | Notes | | |
| |---|---|---|---|---| | |
| | 1 | `audio` | file | β | wav/mp3/flac/ogg + mp4/mov/m4a/webm/β¦ (video & aac-family are ffmpeg-extracted) | | |
| | 2 | `token` | str | `""` | required only when the Space secret `AUDIO_BRIEF_API_TOKEN` is set | | |
| | 3 | `fast` | bool | `false` | skip demucs stems + bass-MIDI β ~2 s instead of ~30 s; prompt omits stem/bassline detail | | |
| | 4 | `bpm_prior` | str | `"default"` | genre slug or numeric string to seed the beat tracker | | |
| | 5 | `llm` | bool | `false` | run the Pollinations LLM to write a natural-language prompt (analyse β brief β lens). Needs `POLLINATIONS_API_KEY` on the Space. **Falls back to the deterministic prompt if the LLM errors** β the call never fails because of the LLM | | |
| | 6 | `lens` | str | `"loose"` | LLM constraint: `match` (tightest, keeps section timings), `loose` (feel + BPM/key, free arrangement), `free` (mood + BPM only). Ignored when `llm=false` | | |
| | 7 | `original_prompt` | str | `""` | the caller's current gen prompt; non-empty switches both paths to anchored steering (measured tempo/key locked, user vocabulary authoritative) | | |
| | 8 | `embedding` | bool | `false` | also run the CLAP similarity stage and return the L2-normalized vector inline as `embedding` (float list) + `embedding_dim`. Runs as a **subprocess island** (torch/checkpoint never resident next to a gen); off by default β heavy, CPU-bound | | |
| ## Two prompt modes | |
| - **Deterministic (default, `llm=false`)** β the prompt is measurements | |
| formatted as structured text. No wallet, no cost, works for any caller. | |
| - **LLM-polished (`llm=true`)** β analyse β LLM writes a brief (infers | |
| genre/mood from the numbers) β LLM compresses it into a natural-language | |
| SA3 prompt. Costs Pollinations pollen, billed to the Space's | |
| `POLLINATIONS_API_KEY`. If that call fails (no key / out of pollen / | |
| timeout) the response degrades to the deterministic prompt and says so | |
| in `prompt_source`. | |
| ## Returns | |
| ```json | |
| { | |
| "prompt": "Sparse ambient micro-cue in A minor Β· BPM β¦", | |
| "prompt_source": "llm:claude-sonnet-4-6", | |
| "brief": "This is an extremely brief, near-static cue in A minor β¦", | |
| "match_style_prompt": "β¦", | |
| "embedding": [0.031, -0.052, β¦], // 512 floats, L2-normalized; null unless embedding=true | |
| "embedding_dim": 512, // null unless embedding=true | |
| "embedding_ckpt": "default-630k", // which CLAP checkpoint made it; don't compare across differing values | |
| "measured": { | |
| "bpm": 128.0, "key": "A", "key_mode": "minor", "key_confidence": 0.82, | |
| "duration_s": 30.0, "lufs_i": -9.1, "lufs_lra": 4.2, "true_peak_db": -0.3, | |
| "sections": [...], "stems": ["bass","drums","other","vocals"], | |
| "stem_stats": {...}, "voiceover_present": false, "bass_midi": {...} | |
| }, | |
| "fast_mode": false, | |
| "stages_ok": ["bpm_key","sections","loudness","stems","bass_midi"], | |
| "errors": [] | |
| } | |
| ``` | |
| `prompt_source` is `"llm:<model>"` on the LLM path, `"deterministic"` when | |
| `llm=false`, or `"deterministic (llm failed: β¦)"` when an LLM run fell back. | |
| `brief` is `null` unless `llm=true`. | |
| ## Python (recommended) | |
| ```python | |
| from gradio_client import Client, handle_file | |
| c = Client("kalamishere/audio-brief") # or the full .hf.space URL | |
| # deterministic (no cost) | |
| r = c.predict( | |
| handle_file("track.wav"), | |
| "YOUR_TOKEN", # "" if the Space isn't gated | |
| True, # fast=True | |
| "default", | |
| False, # llm=False | |
| "loose", | |
| api_name="/audio_to_prompt", | |
| ) | |
| print(r["prompt"]) | |
| # LLM-polished (Space pays pollen) | |
| r = c.predict( | |
| handle_file("track.wav"), | |
| "YOUR_TOKEN", False, "default", | |
| True, # llm=True | |
| "loose", # lens: match | loose | free | |
| api_name="/audio_to_prompt", | |
| ) | |
| print(r["prompt"], "\n\nvia", r["prompt_source"]) | |
| ``` | |
| ## Enabling the LLM path | |
| Set a second Space secret (Settings β Variables and secrets) so the LLM | |
| calls have a wallet to bill: | |
| ``` | |
| POLLINATIONS_API_KEY = <your Pollinations bearer token from enter.pollinations.ai> | |
| ``` | |
| Without it, `llm=true` calls still succeed β they just fall back to the | |
| deterministic prompt (you'll see the reason in `prompt_source`). | |
| ## Gating (protect the CPU-heavy endpoint) | |
| The public Space runs demucs on every full call, so leave it gated. Set a | |
| Space secret in **Settings β Variables and secrets**: | |
| ``` | |
| AUDIO_BRIEF_API_TOKEN = <a long random string> | |
| ``` | |
| When set, calls must pass a matching `token` (constant-time compared) or | |
| get `invalid or missing API token`. When unset, the endpoint is open β | |
| fine for local runs. | |
| ## β οΈ The LLM path is slow β use `submit()`, not `predict()` | |
| `llm=true` makes **two** Pollinations calls (brief + prompt), so a full | |
| request runs ~40 s. `gradio_client`'s synchronous `predict()` has a short | |
| HTTP read timeout and will raise `httpx.ReadTimeout` on it. Use the | |
| job/poll API instead: | |
| ```python | |
| job = c.submit(handle_file("track.wav"), "YOUR_TOKEN", True, "default", | |
| True, "loose", api_name="/audio_to_prompt") | |
| while not job.done(): | |
| time.sleep(2) | |
| r = job.result() # blocks until the queue finishes, no read timeout | |
| print(r["prompt"], "via", r["prompt_source"]) | |
| ``` | |
| Deterministic calls (`llm=false`, `fast=true`) finish in ~2 s and are fine | |
| with plain `predict()`. If you want the LLM path to fit a single `predict()` | |
| call, it can be collapsed to one Pollinations request (skip the separate | |
| brief) β ~20 s and half the pollen, slightly less mood nuance. | |
| ## Similarity embedding (CLAP) | |
| Pass `embedding=true` to get a **512-float, L2-normalized CLAP vector** back in | |
| `embedding`. Because it's L2-normalized, cosine similarity between two tracks is | |
| just the dot product β that's the distance measurement for "how far this track | |
| sits from a trending cluster." | |
| - **Deterministic on full-length audio.** laion-clap (fusion off) randomly | |
| truncates anything longer than its ~10 s window, so a naive whole-track call | |
| returns a *different* vector every time (cos ~0.4-0.9 between repeats of the | |
| same file). The worker instead tiles the audio into fixed 10 s windows, embeds | |
| each, and means them β same file β same vector (repeat cos 1.0). A self-test | |
| repeat-embed guard fails the run if that ever regresses. | |
| - **`embedding_ckpt`** names the checkpoint that produced the vector. Vectors | |
| from different checkpoints are NOT comparable β gate any distance computation | |
| on matching `embedding_ckpt`. | |
| - **Stream-safe by construction.** CLAP is loaded only in `clap_worker.py`, a | |
| short-lived subprocess that embeds one file and exits β torch and the | |
| checkpoint are never resident in the API process, so an embed can't OOM or | |
| crash an in-flight SA3 generation. A failure lands as an `errors` entry and | |
| `embedding: null`; the rest of the response still stands. | |
| - **Poisoned-checkpoint guard.** On load the worker embeds sine / white-noise / | |
| silence and asserts they land far apart. A broken checkpoint (everything | |
| collapses to one vector, top-k still *looks* plausible) fails the call loudly | |
| instead of returning a silently-wrong vector. Threshold: `CLAP_SELFTEST_MAXCOS` | |
| (default 0.90). | |
| - **Checkpoint.** Defaults to the auto-downloaded general 630k-audioset | |
| checkpoint (HTSAT-tiny). Set `LAION_CLAP_MUSIC_CKPT` to the | |
| `music_audioset_epoch_15_esc_90.14.pt` path to use the music-specialized | |
| checkpoint (HTSAT-base) β the worker auto-matches the architecture. | |
| - **Cost.** CPU-bound, ~10-15 s on top of the analysis. Leave it off for runs | |
| that only need BPM/key/sections. | |
| ## Notes / limits | |
| - **Speed:** full analysis is demucs-bound (~30 s on HF free CPU). Use | |
| `fast=true` for interactive/high-volume use, or upgrade the Space tier. | |
| - **Concurrency:** the queue caps concurrency at 4 (see `demo.queue`). | |
| Heavy parallel traffic on the free tier will swap β see the low-RAM | |
| ceiling note in the pipeline docs. | |
| - **YouTube isn't an API input** β pass a file. (The UI's YouTube box | |
| yt-dlp's to a file first; that flow is UI-only.) | |