Diffusers
ONNX
Safetensors
Transformers
English
Chinese
audio-text-to-video
audio-image-text-to-video
audio-driven-video-continuation
avatar
video-generation
Instructions to use halooo7866/LongCat-Video-Avatar with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use halooo7866/LongCat-Video-Avatar with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("halooo7866/LongCat-Video-Avatar", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Transformers
How to use halooo7866/LongCat-Video-Avatar with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("halooo7866/LongCat-Video-Avatar", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 809 Bytes
a8c6f0a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 | {
"_class_name": "LongCatVideoAvatarTransformer3DModel",
"_diffusers_version": "0.32.0",
"adaln_tembed_dim": 512,
"bsa_params":{
"sparsity": 0.9375,
"chunk_3d_shape_q": [4, 4, 4],
"chunk_3d_shape_k": [4, 4, 4]
},
"caption_channels": 4096,
"cp_split_hw": null,
"depth": 48,
"enable_bsa": false,
"enable_flashattn3": false,
"enable_flashattn2": true,
"enable_xformers": false,
"frequency_embedding_size": 256,
"hidden_size": 4096,
"in_channels": 16,
"text_tokens_zero_pad": true,
"mlp_ratio": 4,
"num_heads": 32,
"out_channels": 16,
"patch_size": [
1,
2,
2
],
"audio_window": 5,
"intermediate_dim": 512,
"output_dim": 768,
"context_tokens": 32,
"vae_scale": 4,
"audio_prenorm": false,
"class_range": 24,
"class_interval": 4
}
|