Instructions to use SPRINGLab/SPRING_F5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SPRINGLab/SPRING_F5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="SPRINGLab/SPRING_F5", trust_remote_code=True)# Load model directly from transformers import SPRING_F5 model = SPRING_F5.from_pretrained("SPRINGLab/SPRING_F5", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| ## Backbones quick introduction | |
| ### unett.py | |
| - flat unet transformer | |
| - structure same as in e2-tts & voicebox paper except using rotary pos emb | |
| - possible abs pos emb & convnextv2 blocks for embedded text before concat | |
| ### dit.py | |
| - adaln-zero dit | |
| - embedded timestep as condition | |
| - concatted noised_input + masked_cond + embedded_text, linear proj in | |
| - possible abs pos emb & convnextv2 blocks for embedded text before concat | |
| - possible long skip connection (first layer to last layer) | |
| ### mmdit.py | |
| - stable diffusion 3 block structure | |
| - timestep as condition | |
| - left stream: text embedded and applied a abs pos emb | |
| - right stream: masked_cond & noised_input concatted and with same conv pos emb as unett | |