Spaces:
Paused
Paused
A newer version of the Gradio SDK is available: 6.24.0
metadata
title: Omnivoice
emoji: ๐
colorFrom: pink
colorTo: blue
sdk: gradio
sdk_version: 6.22.0
python_version: '3.12'
app_file: app.py
pinned: false
short_description: omnivoice
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Batch generation
One reference audio + one reference text + N target texts (one per line) โ N generated clips, returned in input-line order.
Calling the API
The Space exposes a /batch_generate endpoint. Install the client with
pip install gradio_client.
from gradio_client import Client, handle_file
# Space id, or "http://127.0.0.1:7860/" when running locally.
client = Client("your-username/omnivoice")
paths, status = client.predict(
texts="The first line to generate.\n"
"The second line to generate.\n"
"And a third one.",
ref_audio=handle_file("reference.wav"), # or a URL
ref_text="Transcript of the reference audio.", # "" to auto-transcribe
language="Auto", # or e.g. "English"; applies to every line
batch_size=8, # lines per GPU call
num_step=32,
guidance_scale=2.0,
denoise=True,
speed=1.0, # ignored when duration is set
duration=None, # fixed seconds per line; None = estimate
preprocess_prompt=True,
postprocess_output=True,
api_name="/batch_generate",
)
print(status)
for i, path in enumerate(paths, start=1):
print(i, path) # local .wav downloaded by the client, in input order
paths is a list of downloaded .wav files โ element i corresponds to line
i of texts. On failure paths is empty and status carries the error
message, so check it:
if not paths:
raise RuntimeError(status)
Copy the files somewhere permanent; the client writes them to a temp directory:
import shutil
for i, path in enumerate(paths, start=1):
shutil.copy(path, f"out_{i:03d}.wav")
Notes
- Only
textsandref_audioreally matter; the rest have the defaults shown above.client.view_api()prints the exact signature if you want to confirm. - Leaving
ref_text=""runs ASR on the reference audio once, then reuses that transcript for every line. Passing the real transcript is faster and more accurate. - Blank lines in
textsare skipped. The maximum is 200 lines per call. - On ZeroGPU each sub-batch is a separate GPU request, so a long list draws
down your quota. Raise
batch_sizefor fewer, larger requests; lower it if you hit out-of-memory errors.