zen-foley / README.md

zeekay

Update model card: add zen/zenlm tags, fix branding

f7dd322 verified about 2 months ago

preview code

raw

history blame contribute delete

1.22 kB

metadata

language: en
license: apache-2.0
tags:
  - text-to-audio
  - zen
  - zenlm
  - hanzo
  - foley
  - sound-effects
  - audio
pipeline_tag: text-to-audio
library_name: transformers

Zen Foley

Foley sound effects generation model for video and interactive media production.

Overview

Built on Zen MoDE (Mixture of Distilled Experts) architecture with 1B parameters.

Developed by Hanzo AI and the Zoo Labs Foundation.

Quick Start

from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
import torch

model_id = "zenlm/zen-foley"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")

# Load audio
import librosa
audio, sr = librosa.load("audio.wav", sr=16000)
inputs = processor(audio, sampling_rate=sr, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs)
print(processor.batch_decode(outputs, skip_special_tokens=True)[0])

Model Details

Attribute	Value
Parameters	1B
Architecture	Zen MoDE
Context	10s audio
License	Apache 2.0

License

Apache 2.0