--- title: Basyx FFmpeg Rendering Engine emoji: 🎬 colorFrom: green colorTo: yellow sdk: docker app_port: 7860 pinned: false license: mit --- # Basyx FFmpeg Rendering Engine Basyx FFmpeg is a CPU-first rendering backend for short-form video automation. It is designed for Hugging Face Docker Spaces running on 2-8 CPU cores with 8-16 GB RAM and no GPU. The engine uses FFmpeg and FFprobe subprocess pipelines and never loads complete videos into Python memory. ## Architecture ```mermaid flowchart TD API["FastAPI REST API"] --> Jobs["Job Manager"] Dashboard["Gradio Operator Dashboard"] --> Jobs Jobs --> Engine["Render Engine"] Engine --> Assets["Asset Probe + Metadata Cache"] Engine --> Ingest["Remote URL + Upload Ingestion"] Engine --> Normalize["Normalization"] Engine --> Scenes["Scene Timeline"] Engine --> Subtitles["SRT / ASS Subtitle Engine"] Engine --> Transitions["Transition Filter Builder"] Engine --> Audio["Voiceover + Ducking Mixer"] Engine --> Exports["Export Manager"] Ingest --> FFmpeg["FFmpeg / FFprobe"] Normalize --> FFmpeg Scenes --> FFmpeg Subtitles --> FFmpeg Transitions --> FFmpeg Audio --> FFmpeg ``` Package layout: ```text renderer/ core/ settings, models, ingestion, orchestration ffmpeg/ command builder, runner, probing, normalization subtitles/ SRT and ASS generation transitions/ xfade graph generation audio/ voiceover and music ducking scenes/ timeline validation templates/ caption templates exports/ final deliverables jobs/ durable job records and retries ``` ## REST API All API endpoints are served at the Space root. Asset fields may be absolute container paths or public `http`/`https` URLs. Remote assets are downloaded into a job-local temp directory with timeout and size limits before rendering. ### `POST /render` Submit one render job. ```json { "template": "tiktok_classic", "output_name": "campaign_clip.mp4", "voiceover": "https://cdn.example.com/voiceover.wav", "background_music": "https://cdn.example.com/music.mp3", "subtitle_format": "ass", "auto_subtitles": true, "subtitle_language": "en", "whisper_model_size": "tiny", "preset": "tiktok_9_16_fast", "callback_url": "https://n8n.example.com/webhook/render-complete", "export_target": "local", "audio_normalize": true, "preview": false, "normalize": true, "scenes": [ { "start": 0, "duration": 5, "media": "https://cdn.example.com/scene1.mp4", "caption": "Launch faster with automated rendering", "transition": "fade" } ] } ``` Response: ```json { "job_id": "job_abc123", "status_url": "/status/job_abc123", "download_url": "/download/job_abc123?token=..." } ``` ### `GET /presets` Returns available render presets and caption templates. Presets let n8n submit compact requests such as `tiktok_9_16_fast`, `youtube_shorts_hd`, `podcast_square`, `reels_with_subtitles`, and `draft_preview`. ### `POST /render/ai-reels` Submit a single-call AI Reels job using an existing voiceover and asset list. TTS is intentionally provider-pluggable in v1; the production path requires a supplied voiceover. ```json { "script": "Launch faster with automated rendering.", "voiceover": "https://cdn.example.com/voiceover.wav", "assets": ["https://cdn.example.com/scene1.jpg", "https://cdn.example.com/scene2.mp4"], "template": "youtube_shorts", "output_name": "ai_reel.mp4" } ``` ### `POST /render/batch` Submit multiple render jobs. The worker pool defaults to one active render to avoid RAM exhaustion. ```json { "jobs": [ { "template": "modern_minimal", "output_name": "clip_a.mp4", "scenes": [{"start": 0, "duration": 3, "media": "https://cdn.example.com/a.mp4"}] } ] } ``` ### `POST /scene-builder` Build a timeline from a script and asset list. ```json { "script": "One script can be split across several short scenes.", "assets": ["https://cdn.example.com/a.mp4", "https://cdn.example.com/b.jpg"], "duration": 8, "transition": "fade" } ``` ### `POST /assets/upload` Upload files from automation tools such as n8n and receive staged paths that can be used in `/render`. Multipart form field: - `files`: one or more images, videos, GIFs, or audio files. Response: ```json { "assets": [ { "filename": "clip.mp4", "path": "/app/temp/uploads/abc/clip.mp4", "reference": "upload://clip.mp4" } ] } ``` ### `POST /render/upload` Submit a render job and files in one multipart request. Multipart fields: - `request_json`: render payload JSON. - `files`: uploaded files referenced by `upload://filename`. Example `request_json`: ```json { "template": "tiktok_classic", "output_name": "n8n_clip.mp4", "scenes": [ { "start": 0, "duration": 5, "media": "upload://clip.mp4", "caption": "Rendered from n8n" } ], "voiceover": "upload://voice.wav" } ``` ### `POST /render/ai-reels/upload` Multipart AI Reels endpoint. Use `upload://filename` inside `voiceover`, `background_music`, and `assets`. ### `POST /transcribe` Transcribe local or remote audio/video with faster-whisper. ```json { "audio": "https://cdn.example.com/voiceover.wav", "model_size": "tiny", "language": "en", "task": "transcribe", "word_timestamps": true, "vad_filter": true } ``` Response includes `text`, detected `language`, segment timings, and optional word timings. Use this endpoint from n8n when you want subtitles before rendering. ### `POST /transcribe/upload` Multipart faster-whisper endpoint for direct n8n file uploads. Multipart fields: - `file`: audio or video file. - `model_size`: optional Whisper model, default `WHISPER_MODEL_SIZE`. - `language`: optional ISO language code. - `word_timestamps`: optional boolean, default `true`. ### `POST /subtitles` Generate SRT or ASS files from timed events. ```json { "format": "ass", "template": "tiktok_zoom", "events": [ {"start": 0, "end": 1.2, "text": "Hello world"} ] } ``` ### `GET /status/{job_id}` Returns job state, logs, FFmpeg commands, metrics, output path, and failure reason. States: `PENDING`, `RUNNING`, `FAILED`, `COMPLETED`. ### `GET /download/{job_id}` Returns the completed MP4 deliverable. Downloads require the signed token returned by render submission. If the job is still running, the endpoint returns `409`. ### `POST /cancel/{job_id}` Requests job cancellation. Pending jobs are marked `CANCELLED`; running jobs are marked `CANCEL_REQUESTED` and stop before the next guarded execution point. ### `POST /admin/cleanup` Deletes expired job records, old exports, and staged upload folders. Default retention is controlled by `JOB_RETENTION_SECONDS`. ### `POST /inspect` Probe one local asset path and return MIME type, duration, codecs, bitrate, resolution, FPS, stream list, and cache metadata. ### n8n automation endpoints - `POST /media/inspect` safely downloads a local or remote asset and returns FFprobe metadata as an asynchronous JSON artifact. - `POST /render/template` substitutes validated JSON variables into a named or inline render template. - `POST /render/variants` creates up to 12 platform outputs and returns a render group ID. - `GET /render/groups/{group_id}` returns aggregate and per-variant status. - `POST /toolkit` supports `loudness_analyze`, `silence_detect`, `black_detect`, `scene_detect`, `contact_sheet`, and `hls` tasks. ## n8n Workflow Integration Recommended URL-based flow: 1. Use an HTTP Request node with `POST https://YOUR-SPACE.hf.space/render`. 2. Pass public file URLs from S3, Supabase Storage, Cloudinary, Google Drive direct-download links, or another CDN in `scenes[].media`, `voiceover`, and `background_music`. 3. Set `auto_subtitles` to `true` when you want faster-whisper captions generated from `voiceover`. 4. Add `callback_url` to receive `render.completed`, `render.failed`, or `render.cancelled` events. Toolkit and analysis jobs use the corresponding `task.*` event names. 5. Poll `GET https://YOUR-SPACE.hf.space/status/{{$json.job_id}}` only if you do not use callbacks. 6. Download the final MP4 from the signed `download_url`. Multipart flow: 1. Use an HTTP Request node set to multipart form-data. 2. Send `request_json` as a text field. 3. Attach files under the `files` field. 4. Reference those files in JSON as `upload://exact-filename.ext`. Security defaults: - Only `http` and `https` remote assets are accepted. - Private, localhost, link-local, and multicast hosts are blocked by default. - Set `ALLOW_PRIVATE_ASSET_URLS=true` only for trusted self-hosted deployments. - `MAX_DOWNLOAD_BYTES` limits remote downloads and multipart upload totals. - `WHISPER_MODEL_SIZE=tiny` and `WHISPER_COMPUTE_TYPE=int8` are the recommended CPU defaults. - `BASYX_SIGNING_SECRET` should be set as a Space secret before public deployment. ## Operator Dashboard The dashboard is available at `/dashboard`. It provides production operator tools for render submission, batch submission, AI Reels submission, job status, logs, downloads, and asset inspection. ## Caption Templates Built-in templates: - `tiktok_classic` - `tiktok_zoom` - `alex_hormozi` - `modern_minimal` - `youtube_shorts` - `podcast_style` - `news_style` Templates are selected per request and do not require code changes. ## CPU And RAM Tuning - Keep `MAX_RENDER_WORKERS=1` for 8 GB RAM deployments. - Use `OUTPUT_PRESET=veryfast` or `ultrafast` for faster CPU rendering. - Use `OUTPUT_CRF=23-28` to balance quality and file size. - Keep source assets near the target duration to reduce normalization work. - Prefer pre-trimmed voiceovers and assets for batch workloads. - Start with `WHISPER_MODEL_SIZE=tiny` or `base` on free/CPU Spaces. - Use `preview=true` for low-resolution draft renders. - Use `audio_normalize=true` for voiceover or podcast-style content. ## Reliability - FFmpeg subprocesses are killed after `FFMPEG_TIMEOUT_SECONDS`. - Remote asset downloads are killed after `DOWNLOAD_TIMEOUT_SECONDS`. - Remote and multipart input size is capped by `MAX_DOWNLOAD_BYTES`. - faster-whisper models are loaded lazily and cached under `WHISPER_MODEL_DIR`. - Webhook callbacks retry with exponential backoff. `CALLBACK_MAX_RETRIES` and `CALLBACK_RETRY_SECONDS` control delivery. - Webhooks include `X-Ava2lon-Event` and an HMAC `X-Ava2lon-Signature` computed with the signing secret. - Set `PUBLIC_BASE_URL` so submissions, status payloads, and callbacks contain absolute URLs. - `export_target=local` copies completed renders to `STORAGE_DIR`. - `export_target=https://...` uploads the completed MP4 with HTTP `PUT`, which works with presigned URLs from S3, Supabase Storage, and similar providers. - Jobs retry up to `MAX_RETRIES`. - Intermediate files are created under `TEMP_DIR` and removed after export. - Job records retain command history, logs, render time, output size, SHA-256 artifact checksums, callback delivery state, and failure reasons. ## Docker Deployment Build locally: ```bash docker build -t basyx-ffmpeg . docker run --rm -p 7860:7860 basyx-ffmpeg ``` Open: - API: `http://localhost:7860/health` - Dashboard: `http://localhost:7860/dashboard` ## Hugging Face Spaces Deployment Create a Docker Space, push this repository, and keep the README metadata above. The container listens on port `7860`, and the Space exposes the FastAPI service plus the dashboard. ## Testing Run: ```bash pytest --cov=renderer --cov-report=term-missing ``` Integration tests use mocked FFmpeg/FFprobe where possible and small generated media where necessary. ## Known v1 Limits - Bundled TTS is not included. The AI Reels endpoint requires a supplied voiceover and keeps narration generation behind a provider interface for future integration. - Advanced caption animation is implemented through ASS effects and FFmpeg-compatible filter behavior, not GPU animation layers. - Batch rendering is intentionally sequential by default for constrained CPU/RAM Spaces.