marquee / README.md
mamuflih13's picture
Remove outdated debugging spots from README for clarity and relevance
d532d78
|
Raw
History Blame
3.96 kB
metadata
title: Marquee
emoji: πŸŽ™οΈ
colorFrom: purple
colorTo: pink
sdk: gradio
sdk_version: 6.18.0
app_file: app.py
license: apache-2.0
short_description: Turn any clip into a broadcast starring your friends
tags:
  - build-small-hackathon

Marquee πŸŽ™οΈ β€” Everyone's a headliner.

Drop a clip of friends doing anything. Marquee recognizes who's in it, you give each person a stage name, pick a show style, and an AI commentator calls the action play-by-play, by name, in sync β€” as a broadcast you can play and export.

βš™οΈ Running it on a Space (read this first)

  • Hardware must be ZeroGPU. The commentary model runs inside @spaces.GPU, which only has a GPU on Nvidia ZeroGPU hardware. On a plain CPU Space the 7B VLM has no GPU and is effectively unusable. Set it in Settings β†’ Hardware β†’ ZeroGPU.
  • No secrets needed at runtime. Everything downloads from public sources (Qwen from the HF Hub, the face models from Intel's OMZ storage). Do not commit any HF_TOKEN/.env β€” if you need one for gradio deploy, keep it local and add it as a Space secret, never a file in the repo.
  • Nothing to pre-download. download_models() fetches the three tiny OpenVINO IR files at first boot (~5 MB total). First boot also warms up Qwen.

How it's wired (the important part)

The backend is a gradio.Server (a FastAPI subclass β€” Gradio 6). It gives us ZeroGPU + the gradio runtime while letting our own routes take priority. The custom UI is a static HTML page at /. The CPU-only scan is a plain fetch() route; the GPU commentary is an @app.api() endpoint called through the Gradio JS client, which is what wires it into the queue and lets ZeroGPU detect the @spaces.GPU function. No iframe, no SSR fighting us for /.

Static UI (Marquee.html + marquee.css + marquee.js)  ──►  served at "/"
   β”‚  fetch() for scan Β· Gradio JS client for generate
   β–Ό
FastAPI (app.py)
   β”œβ”€ POST /api/scan      (file)                -> normalize 720p -> detect+cluster faces
   β”‚                                               -> session + roster (face crops as data URLs)
   β”œβ”€ GET  /video/{sid}                         -> streams the normalized mp4
   └─ @app.api /generate  (session, names, vibe)-> key events -> name-grounded frames
                                                   -> Qwen2.5-VL (ZeroGPU) -> {t,text} script

Heavy per-session state (identities, motion, normalized video path) lives in a server-side SESSIONS dict keyed by the session_id the scan call returns, so the two stateless client calls never round-trip embeddings.

Face recognition is used only to feed names to the VLM β€” boxes are never shown. The broadcast player reads the real normalized <video> and syncs the chyron + play-by-play rail to video.currentTime.

Files

  • app.py β€” FastAPI backend + the two API endpoints + UI assembly
  • Marquee.html / marquee.css / marquee.js β€” the custom UI (design + flow)
  • video.py β€” rotation-aware 720p normalize (one upright source of truth)
  • faces.py β€” OpenVINO detect β†’ align β†’ embed β†’ cluster (CPU)
  • events.py β€” motion-based key events
  • commentary.py β€” Qwen2.5-VL personas (incl. Diva Hour) + strict JSON parse
  • ov_models.py β€” OpenVINO IR auto-download + thin runtime wrappers

Models (all open)

  • face-detection-retail-0004, landmarks-regression-retail-0009, face-reidentification-retail-0095 (OpenVINO IR, CPU)
  • Qwen2.5-VL-7B-Instruct (ZeroGPU)

North star (next iteration)

  • TTS per line (hyped delivery) mixed under the clip's own audio.
  • Rendered MP4 export with captions (and voice) burned in.
  • Canvas hype FX on key moments (data's already there).

Local dev

pip install -r requirements.txt
python -c "from ov_models import download_models; download_models()"
python app.py    # http://localhost:7860  (Qwen lazy-loads on first generate)