--- license: apache-2.0 tags: - executorch - xnnpack - pte - on-device - audio-classification - zero-shot-classification - feature-extraction base_model: - laion/clap-htsat-fused --- # CLAP HTSAT-fused — ExecuTorch (search sound with words) CLIP for audio. An audio tower and a text tower put a clip and a phrase in the same 512-dimensional space, so "a dog barking" can be matched against a recording with a cosine — with words the app chooses at runtime, not a label list fixed at training time. The shelf already has AST, which answers with AudioSet's 527 labels, and CLIP and SigLIP2, which do this for pictures. This is the piece that was missing: a searchable representation of sound. ``` audio: input_features (1,4,1001,64) -> (1,512) text : (input_ids (1,32), attention_mask (1,32)) -> (1,512) ``` Both embeddings come out L2-normalised, so a cosine is a dot product. Two files because the text side runs once per phrase and can be cached; the audio side runs once per clip. | tower | build | file | MB | embedding cosine vs eager | Mac ms* | |---|---|---|---|---|---| | audio | XNNPACK fp32 | `clap_audio_xnnpack_fp32.pte` | 113.1 | 1.000000 | 115.5 | | audio | XNNPACK int8 | `clap_audio_xnnpack_int8.pte` | **29.1** | 0.999613 | 115.2 | | text | XNNPACK fp32 | `clap_text_xnnpack_fp32.pte` | 501.4 | 1.000000 | 11.6 | | text | XNNPACK int8 | `clap_text_xnnpack_int8.pte` | **243.2** | 0.990783 | 12.4 | | text | Core ML (iOS) | `clap_text_coreml_all.pte` | 251.2 | 0.999999 | 2.5 | The int8 pair is **272 MB** against fp32's 615 MB, and both keep all five winning phrases. The audio tower quantises well (0.9996); the text tower is the one to watch at 0.9908, which clears this shelf's 0.99 bar for embeddings but not by much. \*Mac arm64, single process, median of 5 — a reference point for relative cost, not a device number. Torch eager fp32 on the same machine: audio 29.0 ms, text 13.9 ms. XNNPACK delegate coverage: audio 64.2%, text 77.3%. The text tower is the big one — it is a RoBERTa base, 124.65 M parameters against the audio tower's 27.55 M — and it is also the one you can run once and keep. A phrase set of twenty labels is twenty 512-float vectors. **The audio tower has no Core ML build.** coremltools stops on a data-dependent size inside the HTSAT windowing (`Could not extract specialized integer from data-dependent expression u12`). The text tower converts cleanly and is 100% delegated there. ## Running it **1. The audio.** `ClapFeatureExtractor` at 48 kHz mono: 64 mel bins, a 10-second window, and the four-chunk fusion layout the `-fused` checkpoint wants — `input_features` comes out `[1, 4, 1001, 64]`. Feed that straight in. The extractor also returns `is_longer`, and this graph does not take it. That is deliberate, and it is the one thing worth reading before you use this file. **2. The phrase.** The repo's RoBERTa tokenizer with `padding="max_length", max_length=32`. Padding length does not change the embedding — checked, cosine 1.0 against padding to the natural length — but the graph's shape is fixed at 32. **3. The match.** Dot the two vectors. Both are already unit length, so the dot product is the cosine. For a small label set, embed every phrase once, stack them, and take the argmax. ## The `is_longer` flag, and why it is not an input `ClapFeatureExtractor` ends with this: ```python if truncation == "fusion" and sum(is_longer) == 0: # if no audio is longer than 10s, then randomly select one audio to be longer rand_idx = np.random.randint(0, len(input_mel)) is_longer[rand_idx] = True ``` For a batch of one — what an app sends — that means **`is_longer` is always True**, however short the clip. Measured on five macOS system sounds of 0.76 to 1.65 seconds: True every time. So the fusion path is the real path, and the graph takes it unconditionally. The alternative was to keep the `if len(is_longer_idx) > 0` branch, which `torch.export` cannot specialise (`GuardOnDataDependentSymNode ... u0`) because it is a Python branch on the size of a `torch.where` result. Building the other half — global-only, `is_longer` forced False — moves the embedding to **cosine 0.78–0.87** against the stock model and changes which phrase wins. That build is wrong, and the way it first looked right is worth recording: it was compared against a copy of itself that had the same forced flag, and reported `max_abs_diff 0.000e+00`. The control has to move the variable under test. The branch-free fusion graph matches the stock module driven by the extractor's own flag at **max_abs_diff 0.000e+00**. ## What is measured Two numbers, because an embedding model has two things that can go wrong. **Fidelity.** Cosine between each build's embedding and the fp32 model's, over five sounds and six phrases: **1.000000 for both fp32 towers**, and **0.999613 (audio) / 0.990783 (text)** for int8. **Behaviour.** Whether the phrase that wins for each sound is still the same: **5/5 for fp32 and for int8**. ``` Ping an electronic beep Submarine an electronic beep Purr a bubble popping Sosumi an electronic beep Glass a bell ringing ``` ```bash python convert/check_clap.py fp32 # or int8 python convert/audit_int8.py clap_audio ``` The audio is macOS's own system sounds — short real recordings. That matters: the first probe used synthetic speech from `say`, and the **fp32 model** rated it "rain falling" at 0.82 before any conversion was involved. CLAP is trained on environmental audio and speech is not what it is for. Five sounds and six phrases is a conversion check, not an evaluation of the model. - **Source**: [laion/clap-htsat-fused](https://huggingface.co/laion/clap-htsat-fused) - **License**: Apache-2.0 torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))