transcriber / README.md
asuilin's picture
Initial upload
3bb09ca verified
|
Raw
History Blame Contribute Delete
3.04 kB

A newer version of the Gradio SDK is available: 6.21.0

Upgrade
metadata
title: Transcriber
emoji: 🎬
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.19.0
app_file: app.py
pinned: false

transcriber

Download a YouTube video and generate embedded subtitles using ElevenLabs speech-to-text.

What it does

  1. Downloads a YouTube video via yt-dlp (or accepts a local file)
  2. Extracts audio and transcribes it with ElevenLabs (word-level timestamps)
  3. Builds an SRT file using limit-first segmentation with pause detection
  4. Optionally translates subtitles into a second language via LiteLLM
  5. Muxes video + subtitle track(s) into an MKV with soft (toggleable) subtitles

The raw ElevenLabs JSON response is saved alongside the output and reused on reruns, so the API is only called once per video. Translations are cached the same way.

Setup

uv sync

Create a .env file:

ELEVENLABS_API_KEY=your_key_here
VIDEO_LANGUAGE=eng       # optional; ElevenLabs auto-detects if omitted
TRANSLATE_TO=eng         # optional; ISO 639-3 target language for a translated track
TRANSLATION_MODEL=...    # required when TRANSLATE_TO is set; any LiteLLM model string
OUTPUT_DIR=~/subtitled   # optional; default output directory

Requires ffmpeg on PATH.

Usage

Web app

uv run app.py

Opens at http://localhost:7860. Supports the same options as the CLI. The output MKV is available for download directly in the browser and is automatically deleted 30 minutes after the job completes.

CLI

# Download from YouTube and transcribe
uv run main.py "https://www.youtube.com/watch?v=..."

# Use an already-downloaded file
uv run main.py --input-video lecture.mp4

# Common options
uv run main.py "https://..." \
  --output-dir ~/subtitled \
  --language jpn \
  --resolution 1080 \
  --max-length 80 \
  --max-pause 2.0

# Add a translated subtitle track
uv run main.py "https://..." --language jpn --translate-to eng

Options

Option Default Description
--input-video / -i Local video file (skips download)
--output-dir / -o $OUTPUT_DIR or . Directory for all output files
--language / -l $VIDEO_LANGUAGE ISO 639-3 source language code, e.g. eng, jpn
--resolution / -r 720 Target resolution (smallest side, px)
--max-length 100 Max characters per subtitle cue
--max-pause 3.0 Silence gap (s) that triggers a cue split
--translate-to $TRANSLATE_TO ISO 639-3 target language for a translated subtitle track; requires $TRANSLATION_MODEL
--translation-context 3 Number of surrounding cues used as context when translating

Output files

Given a video titled My Lecture with --translate-to eng:

File Description
My Lecture.mp4 Original download (kept)
My Lecture.json Raw ElevenLabs transcription (cache)
My Lecture.eng.json Translated cues cache (only when --translate-to is used)
My Lecture.mkv Video with embedded soft subtitle track(s)