dasheng-base + MAE decoder (AudioSet-Opus 24kbps)

Frozen dasheng-base encoder with a trained MAE decoder (8-layer, d=512), trained for one epoch on danjacobellis/audioset_opus_24kbps to reconstruct masked mel patches (mask ratio 0.75, group factor 2).

The encoder was never updated. This checkpoint is used for hard-example mining: per-clip reconstruction loss is the difficulty signal for SSL data curation.

last.pt / best.pt are PyTorch payloads with model_state (encoder) and objective_state (encoder + decoder). config.json is the full training config.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support