Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

markmuller
/
TTS_POST_trainnig_1_emo

Text-to-Speech
VoxCPM
Safetensors
Persian
English
persian
farsi
emotion
paralinguistic
expressive-tts
Model card Files Files and versions
xet
Community

Instructions to use markmuller/TTS_POST_trainnig_1_emo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • VoxCPM

    How to use markmuller/TTS_POST_trainnig_1_emo with VoxCPM:

    import soundfile as sf
    from voxcpm import VoxCPM
    
    model = VoxCPM.from_pretrained("markmuller/TTS_POST_trainnig_1_emo")
    
    wav = model.generate(
        text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.",
        prompt_wav_path=None,      # optional: path to a prompt speech for voice cloning
        prompt_text=None,          # optional: reference text
        cfg_value=2.0,             # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse
        inference_timesteps=10,   # LocDiT inference timesteps, higher for better result, lower for fast speed
        normalize=True,           # enable external TN tool
        denoise=True,             # enable external Denoise tool
        retry_badcase=True,        # enable retrying mode for some bad cases (unstoppable)
        retry_badcase_max_times=3,  # maximum retrying times
        retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech
    )
    
    sf.write("output.wav", wav, 16000)
    print("saved: output.wav")
  • Notebooks
  • Google Colab
  • Kaggle

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Gated model
You can list files but not access them

Preview of files found in this repository
  • best_checkpoint_of_post_training
    Add best_checkpoint_of_post_training 4 days ago
  • step_0002000
    Add step_0002000 4 days ago
  • step_0006000
    Add step_0006000 4 days ago
  • step_0008000
    Add step_0008000 4 days ago
  • step_0008500
    Add tokenization_voxcpm2.py to step_0008500 4 days ago
  • step_0009000
    Add tokenization_voxcpm2.py to step_0009000 4 days ago
  • step_0009167
    Add tokenization_voxcpm2.py to step_0009167 4 days ago
  • step_0009168
    Add step_0009168 4 days ago
  • .gitattributes
    1.52 kB
    initial commit 4 days ago
  • README.md
    7.29 kB
    Fix usage: per-checkpoint snapshot_download, not whole-repo 4 days ago