STEMMA-MOSS-Music
Paper-final MOSS-Music + STEMMA-Instruct: V7 coreference + relation MultiQA 1x
(no fullsong/block component), frozen audio encoder, global batch 128, one
epoch, LR 1e-5. Final step 245; base revision
fce7f8304e96cc2d3398b8106456cbb2ecec3139. Four BF16 weight shards are at the
repository root and indexed by model.safetensors.index.json.
Status: isolated installation, 17 CPU tests and 24 full-checkpoint B200 requests
passed. Final/DeepStack outputs exactly match the pinned official encoder in
FP32 and BF16 on identical mel/weights. Input tokens match sealed evaluation
records; one parsed answer differs from vLLM in this sample. See VALIDATION.md.
This is not paper score reproduction. This repository contains complete
fine-tuned weights and editable native Transformers/PyTorch inference code,
not a LoRA adapter.
Minimal use and editable features
python -m pip install -r requirements.txt
python examples/generate.py --model . --audio a.wav b.wav
python examples/embeddings.py --model . --audio a.wav b.wav
from stemma import Request, load_model
session = load_model(".")
batch = session.prepare([Request("Audio A: <audio>\nAudio B: <audio>\nCompare.", ["a.wav", "b.wav"])])
features = session.extract_audio_embeddings(batch)
contextual = session.extract_contextual_embeddings(batch, layers=(-1,))
answers = session.generate([Request("Compare these sources.", ["a.wav", "b.wav"])])
The source package includes the pinned MOSS model/processor/config. Use
stemma.load_model, not a bare AutoModel guess or an external base cache. Resolve
and record an explicit HF revision before snapshot_download, then pass the
downloaded local directory to the loader. All required inference assets are local.
from huggingface_hub import HfApi, snapshot_download
import sys
repo = "hoyso48/STEMMA-MOSS-Music"
revision = HfApi().model_info(repo).sha # Record this commit for reproduction.
folder = snapshot_download(repo, revision=revision)
sys.path.insert(0, folder) # The downloaded repository includes the stemma code.
from stemma import load_model
session = load_model(folder)
See API.md and core docstrings for tensor shapes and ordinary forward-hook/grad
access. The encoder uses official_chunked_400_v1: full-source mel followed by
400-frame (~4-second) independent chunks, reset positions, masked padding and
retained short tails. conv_chunksize=64 is a processing batch, not duration.
Final and three DeepStack streams are regrouped into original source/sample order.
The deterministic nonpersistent sinusoid buffer is regenerated. Source time
markers remain separate from chunk positions. Context is 40,960 tokens.
No evaluation audio/dataset, optimizer, RNG or training pickle is distributed.
Read NOTICE.md and licenses/Apache-2.0.txt for upstream terms and local changes;
new wrapper/example code is also Apache-2.0 (LICENSE). Stock vLLM serving is
not supported by this package; the paper's patched backend is described in API.md.
release_info.json, provenance.json and checksums.json identify staged assets.
- Downloads last month
- 18
Model tree for hoyso48/STEMMA-MOSS-Music
Base model
OpenMOSS-Team/MOSS-Music-8B-Instruct