multimodalart's picture
multimodalart HF Staff
Document live backend architecture, add sdk_version 6.12.0
94fce9e verified
|
Raw
History Blame Contribute Delete
1.99 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: ABot-World Interactive
emoji: 🌍
colorFrom: gray
colorTo: red
sdk: gradio
sdk_version: 6.12.0
app_file: app.py
pinned: false
short_description: Live interactive world rollout from an image
python_version: '3.10'
startup_duration_timeout: 40m

🌍 ABot-World β€” Interactive World Rollout

Upload a single starting image and steer a live navigable world in real time. Provide a first-frame image (image-to-video seed), describe the scene, and drive the world with WASD (move & turn) / IJKL (look & pan). The model autoregressively rolls out an action-conditioned world and streams decoded frames straight to your browser.

Architecture

This Space runs a proper live backend/infrastructure (in the spirit of Overworld/waypoint-1-5):

  • gradio.Server exposes ZeroGPU-friendly /start_game and /stop_game API endpoints plus a raw WebSocket (/ws) that streams binary JPEG frames and receives live control input β€” no Gradio UI polling.
  • A custom index.html front-end handles image upload (the i2v seed β€” no video upload needed), live keyboard steering, and a clean HUD.
  • Every endpoint is keyed by a per-client session_id, so concurrent players never share seed images, frame queues, or status messages.
  • ZeroGPU quota is billed to the requesting user: the HF iframe's x-api-token / x-ip-token is forwarded by the Gradio JS client and the incoming request context is propagated into the GPU worker thread, so @spaces.GPU consumes the visitor's own quota rather than the Space owner's.

Inference is causal / block-wise, matching the upstream CausalInferencePipeline.