LongCat-Video-Avatar 1.5
ONNX
Diffusers
Safetensors
Transformers
English
Chinese
audio-text-to-video
audio-image-text-to-video
audio-driven-video-continuation
avatar
video-generation
Instructions to use youngexlance/Vatarstilfly with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LongCat-Video-Avatar 1.5
How to use youngexlance/Vatarstilfly with LongCat-Video-Avatar 1.5:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Diffusers
How to use youngexlance/Vatarstilfly with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("youngexlance/Vatarstilfly", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Transformers
How to use youngexlance/Vatarstilfly with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("youngexlance/Vatarstilfly", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 853 Bytes
643fb20 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 | {
"_class_name": "LongCatVideoAvatarTransformer3DModel",
"architectures": [
"LongCatVideoAvatarTransformer3DModel"
],
"_diffusers_version": "0.32.0",
"in_channels": 16,
"out_channels": 16,
"hidden_size": 4096,
"depth": 48,
"num_heads": 32,
"caption_channels": 4096,
"model_max_length": 512,
"mlp_ratio": 4,
"adaln_tembed_dim": 512,
"frequency_embedding_size": 256,
"patch_size": [
1,
2,
2
],
"enable_flashattn3": false,
"enable_flashattn2": true,
"enable_xformers": false,
"enable_bsa": false,
"bsa_params": null,
"cp_split_hw": null,
"text_tokens_zero_pad": true,
"audio_window": 5,
"audio_block": 5,
"audio_channel": 1280,
"intermediate_dim": 512,
"output_dim": 768,
"context_tokens": 32,
"vae_scale": 4,
"audio_prenorm": false,
"class_range": 24,
"class_interval": 4
} |