sammoran-phd's picture
Add files using upload-large-folder tool
80ac38b verified
|
Raw
History Blame Contribute Delete
3.43 kB
---
license: other
license_name: stability-ai-community-license
license_link: https://huggingface.co/stabilityai/stable-audio-open-small/blob/main/LICENSE
library_name: stable-audio-tools
pipeline_tag: text-to-audio
base_model: stabilityai/stable-audio-open-small
base_model_relation: finetune
tags:
- stable-audio
- music
- cara
- attribution
- research
---
# CARA-native Stable Audio Open Small
This is the Phase 2 CARA-native Stable Audio Open Small checkpoint used in the
CARA cross-architecture attribution research. The fork exposes batch-aligned
native DiT features and trains a checkpoint-owned hierarchical 98-pool/9-family
head together with the diffusion model.
The release is an open-weight peer-review artifact, not an OSI open-source
model. Its use and redistribution remain subject to the Stability AI Community
License.
## Contents
- `phase2/native_stable_audio_full.ckpt`: completed 7,665-step model/head
checkpoint.
- `phase2/native_stable_audio_training_report.json`: training contract.
- `registry/`: exact CARA pool/family ordering used by the head.
- `evidence/`: authoritative prompt-visible and fixed-audio score reports.
- `source/`: exact source snapshot used to load and evaluate the checkpoint.
- `LICENSE` and `NOTICE`: required upstream license and attribution.
- `cara_model_manifest.json`: byte sizes and SHA-256 values for release files.
Use the immutable `phase2-v1` tag, or the Hub commit hash it resolves to.
## Reload
```bash
hf download sammoran-phd/cara-native-stable-audio \
--revision phase2-v1 \
--local-dir cara-native-stable-audio-release
mkdir cara-native-stable-audio-source
tar -xzf \
cara-native-stable-audio-release/source/cara-native-stable-audio-source.tar.gz \
-C cara-native-stable-audio-source
cd cara-native-stable-audio-source
```
Load `stabilityai/stable-audio-open-small` with `stable-audio-tools`, attach the
included fork's `CARAAttributionHead`, and load
`phase2/native_stable_audio_full.ckpt`. The included benchmark scripts perform
the registry-hash, global-step, native-head, and feature-shape checks before
evaluation. The complete invocation is in the `cara-native-musicmodels` Phase 2
job specification.
## Evaluation boundary
On the primary 780-waveform balanced fixed-audio core, this checkpoint scored
7.82% exact top-1, 25.64% top-3, 33.21% pool-derived family accuracy, 6.90% ECE,
and 100% registry-valid output. On the earlier 320-row prompt-visible benchmark
it scored 23.75% exact top-1, 56.56% top-3, and 99.38% family accuracy.
The fixed-audio result is the primary estimate of audio-conditioned attribution.
The much higher prompt-visible family score is dominated by visible semantic
taxonomy and is not evidence of reliable exact source identification.
## Intended use and limitations
This release is intended for academic reproduction, interface auditing, and
controlled attribution experiments. It is not a provenance, copyright
identification, royalty allocation, or safety system. The primary fixed-audio
core covers 39 of 98 pools and represents one source corpus and one training
run.
## License
The `stable-audio-tools` code is MIT licensed. The base model and this derivative
checkpoint are governed by the Stability AI Community License. Redistribution
must include that agreement and the required `NOTICE`; commercial conditions
depend on the user's circumstances and the current upstream license.