lf2ar-speech-2b / README.md
tiagoCuervo's picture
Publish model checkpoint
f164509 verified
|
Raw
History Blame Contribute Delete
1.52 kB
---
license: apache-2.0
library_name: interleaved-lm
tags:
- speech-language-model
- lf2ar
- safetensors
---
# lf2ar-speech-2b
Speech LF²AR checkpoint from *LF²AR: Accounting for Layerwise Dynamics to Improve
Multimodal Adaptation of Language Models*. This model has
1,980,911,664 parameters and uses dynamic attention residuals.
```python
from interleaved_lm import PerceptionExpressionAdaptedTextLM
from interleaved_lm.audio import generate_speech
model = PerceptionExpressionAdaptedTextLM.from_pretrained(
"tiagoCuervo/lf2ar-speech-2b", device="cuda"
)
codes = generate_speech(model, "A short story about the moon", max_new_tokens=250)
```
Speech is represented by consecutive-run-collapsed 25 Hz mHuBERT layer-11 K-means
units (K=500; EOS=500; PAD=501). A vocoder is not bundled.
The text backbone is initialized from
[SmolLM](https://huggingface.co/collections/HuggingFaceTB/smollm-6695016cad7167254ce15966).
Speech-tokenizer and vocoder artifacts are credited in the
[Interleaved-LM third-party notices](https://github.com/tiagoCuervo/interleaved-lm/blob/main/THIRD_PARTY.md).
## Evaluation
| Task | Direction | Accuracy |
|---|---|---:|
| sstorycloze | audio-audio | 61.57 |
| sstorycloze | text-text | 73.70 |
| sstorycloze | audio-text | 63.98 |
| sstorycloze | text-audio | 64.03 |
| tstorycloze | audio-audio | 87.60 |
| tstorycloze | text-text | 92.52 |
| tstorycloze | audio-text | 86.91 |
| tstorycloze | text-audio | 83.97 |
See `provenance.json` for exact hashes and evaluation metadata.