lf2ar-speech-2b

Speech LF²AR checkpoint from LF²AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models. This model has 1,980,911,664 parameters and uses dynamic attention residuals.

from interleaved_lm import PerceptionExpressionAdaptedTextLM
from interleaved_lm.audio import generate_speech

model = PerceptionExpressionAdaptedTextLM.from_pretrained(
    "tiagoCuervo/lf2ar-speech-2b", device="cuda"
)
codes = generate_speech(model, "A short story about the moon", max_new_tokens=250)

Speech is represented by consecutive-run-collapsed 25 Hz mHuBERT layer-11 K-means units (K=500; EOS=500; PAD=501). A vocoder is not bundled.

The text backbone is initialized from SmolLM. Speech-tokenizer and vocoder artifacts are credited in the Interleaved-LM third-party notices.

Evaluation

Task Direction Accuracy
sstorycloze audio-audio 61.57
sstorycloze text-text 73.70
sstorycloze audio-text 63.98
sstorycloze text-audio 64.03
tstorycloze audio-audio 87.60
tstorycloze text-text 92.52
tstorycloze audio-text 86.91
tstorycloze text-audio 83.97

See provenance.json for exact hashes and evaluation metadata.

Downloads last month
7
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support