avatar-demo / README.md
Mike0021's picture
Build unified talking avatar Space
a1210e3 verified
|
Raw
History Blame Contribute Delete
1.38 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: Avatar Demo
emoji: 🎬
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 5.49.1
python_version: '3.10'
app_file: app.py
hardware: zero-a10g
suggested_hardware: zero-a10g
startup_duration_timeout: 1h
license: apache-2.0
short_description: Text, image, and voice to talking avatar video
models:
  - nineninesix/gepard-1.0
  - nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
  - microsoft/wavlm-base-plus
  - meituan-longcat/LongCat-Video
  - meituan-longcat/LongCat-Video-Avatar-1.5
tags:
  - text-to-speech
  - talking-avatar
  - video-generation
  - zerogpu

Avatar Demo

This Space merges GEPARD TTS with LongCat-Video-Avatar 1.5 in one Gradio app.

Workflow:

  1. Enter text.
  2. Upload a reference face image.
  3. Pick a GEPARD preset voice or upload a voice reference clip.
  4. Generate speech audio and a 5-second lip-synced avatar video.

The GEPARD checkpoint is private, so the Space must have an HF_TOKEN secret with access to nineninesix/gepard-1.0.

ZeroGPU notes:

  • create_env.py re-pins transformers==5.3.0 at startup before ML imports.
  • spaces is imported before torch.
  • Both model stacks are loaded at module level and moved to cuda.
  • The generation handler is wrapped with @spaces.GPU(..., size="xlarge") using a duration callable that returns 300 seconds.
  • LongCat's xformers call sites use a local PyTorch SDPA shim.