LongCat-Video-Avatar 1.5
ONNX
Diffusers
Safetensors
Transformers
English
Chinese
audio-text-to-video
audio-image-text-to-video
audio-driven-video-continuation
avatar
video-generation
Instructions to use smartdigitalnetworks/LongCat-Video-Avatar-1.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LongCat-Video-Avatar 1.5
How to use smartdigitalnetworks/LongCat-Video-Avatar-1.5 with LongCat-Video-Avatar 1.5:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Diffusers
How to use smartdigitalnetworks/LongCat-Video-Avatar-1.5 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("smartdigitalnetworks/LongCat-Video-Avatar-1.5", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Transformers
How to use smartdigitalnetworks/LongCat-Video-Avatar-1.5 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("smartdigitalnetworks/LongCat-Video-Avatar-1.5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| { | |
| "_class_name": "LongCatVideoAvatarTransformer3DModel", | |
| "architectures": [ | |
| "LongCatVideoAvatarTransformer3DModel" | |
| ], | |
| "_diffusers_version": "0.32.0", | |
| "in_channels": 16, | |
| "out_channels": 16, | |
| "hidden_size": 4096, | |
| "depth": 48, | |
| "num_heads": 32, | |
| "caption_channels": 4096, | |
| "model_max_length": 512, | |
| "mlp_ratio": 4, | |
| "adaln_tembed_dim": 512, | |
| "frequency_embedding_size": 256, | |
| "patch_size": [ | |
| 1, | |
| 2, | |
| 2 | |
| ], | |
| "enable_flashattn3": false, | |
| "enable_flashattn2": true, | |
| "enable_xformers": false, | |
| "enable_bsa": false, | |
| "bsa_params": null, | |
| "cp_split_hw": null, | |
| "text_tokens_zero_pad": true, | |
| "audio_window": 5, | |
| "audio_block": 5, | |
| "audio_channel": 1280, | |
| "intermediate_dim": 512, | |
| "output_dim": 768, | |
| "context_tokens": 32, | |
| "vae_scale": 4, | |
| "audio_prenorm": false, | |
| "class_range": 24, | |
| "class_interval": 4 | |
| } |