A newer version of the Gradio SDK is available: 6.22.0
Immersive AI Time Machine β Research & Implementation Plan
The Vision
Transform the AI Time Machine from a voice-chat-with-text-panels into a portal experience β when you click "Launch," the screen dissolves into the destination world, a character appears before you, and you converse face-to-face while the ambient sounds of that era surround you.
Current State Assessment
Your app currently renders everything as Gradio markdown panels. The Destination model already has rich data (atmosphere, visual_motifs, visual_preset_key) that's never used visually β it's just displayed as text. The Persona has speaking_style, voice_theory, etc., but there's no visual character representation. You already have great bones β the steampunk cockpit CSS, the live voice WebSocket channel, the streaming TTS. What's missing is the sensory layer.
Current: [Text Description] β [Voice Chat] β [Text Transcript]
Target: [360Β° Scene + Ambient Audio] β [Talking Avatar] β [Immersive UI]
User Review Required
This plan is organized into 3 tiers β each tier is independently shippable and adds progressively more immersion. You can ship Tier 1 alone for a massive upgrade, then layer on Tier 2 and 3 later. Which tier(s) do you want to pursue?
Hackathon budget considerations: Some of these tools have free tiers, some don't. I've flagged costs for each. The AI image generation approach (Tier 1) is the most cost-effective since you already have a Together AI API key. 3D avatars (Tier 2+) add monthly API costs.
Tier 1: Visual World & Character (Highest Impact, Lowest Complexity)
This tier alone transforms the experience dramatically. No new backend services needed β just AI-generated images + frontend magic.
1A. AI-Generated Scene Backgrounds
What: When a Destination is generated, use an image generation API to create a matching scene that fills the screen.
| Tool | How It Works | Pricing | Best For |
|---|---|---|---|
| Blockade Labs Skybox AI β | Text β 360Β° equirectangular panorama + 3D mesh. Up to 8K resolution | ~$20/mo (100 credits, 3-6 per gen) | Full 360Β° immersion in Three.js |
| Flux 2 (via Together AI) β | Text β image via the same Together AI you're already using for LLM | ~$0.02-0.05/image, same API key | Zero new API signup |
| Google Imagen 4 | Text β high-quality scene image via Vertex AI | $0.02-0.06/image ($300 GCP trial) | High quality, Google ecosystem |
| Stability AI (SD 3.5) | Text β scene image. Open-source models available for self-hosting | 25 free credits. ~$0.01-0.08/image | Self-hosting option for unlimited free gen |
| OpenAI GPT Image 2 | Text β image. Exceptional prompt adherence | $0.005-0.21/image | Complex scene descriptions |
Recommended: Flux 2 via Together AI for scene backgrounds. You're already using Together AI for your LLM (
TIME_MACHINE_LLM_API_KEY). Their image API supports Flux models with the same API key. Zero friction. For maximum wow, Blockade Labs gives you true 360Β° panoramas that users can look around in.
Implementation approach:
- After
generate_destination(), call image API with a prompt built fromdestination.atmosphere+destination.place+destination.year - Return image URL/bytes as part of the
DestinationEvent - Frontend: fade-in the scene as a full-bleed background behind the UI, with a vignette overlay
1B. AI-Generated Character Portraits
What: When a Persona is generated, create a portrait of that character.
| Tool | Approach | Key Advantage |
|---|---|---|
| Flux 2 Kontext β | Generate a portrait from persona description | Best for character consistency β can maintain features across multiple generations using reference images |
| Imagen 4 | Generate from persona + era description | Deep Google ecosystem integration |
| Stability AI | Generate + can self-host for free | Open-source, most flexible |
Flux 2 Kontext is specifically designed for maintaining character consistency across multiple generations β essential if you want the same character to look the same across scene changes or re-generations.
Implementation approach:
- After
generate_persona(), call image API with a prompt built from persona description + destination era - Display the character portrait in a circular "aperture" frame (you already have the
.tm-cockpit-apertureCSS!) - During conversation, the portrait is always visible β this is who you're talking to
1C. Immersive Frontend Overhaul
What: Replace the current Gradio markdown panels with a cinematic, full-screen experience.
| Effect | How | Impact |
|---|---|---|
| Full-bleed scene background | Scene image as background-image with background-size: cover, dark vignette overlay |
Instant world transportation |
| Cinematic fade-in | GSAP-powered transition from cockpit β wormhole β destination | Transition magic |
| Parallax depth | Scene background moves slightly on mouse movement, character portrait floats above | Depth and life |
| Atmospheric particles | Canvas-based floating particles (embers for fire scenes, snow for winter, dust for ancient ruins) | Living world |
| Film grain + color grading | CSS filter + animated noise overlay with mix-blend-mode: overlay (opacity 0.03-0.08) |
Era authenticity |
| Era-specific color grading | Sepia for ancient, blue-tint for future, warm amber for candlelit eras | Visual time-travel |
| Ambient glow pulse | Character portrait pulses with soft glow when TTS audio is playing | Character feels alive |
| Floating dialogue text | Conversation text fades in/out near the character, RPG-dialogue style | Immersive dialogue |
GSAP (GreenSock Animation Platform) is now 100% FREE β including commercial use β after its Webflow acquisition. It's the industry gold standard for web animation timing and sequencing. Perfect for cinematic launch transitions.
1D. Ambient Soundscapes
What: Background audio that matches the destination (marketplace chatter, ocean waves, wind, etc.)
| Tool | How | Pricing | Key Feature |
|---|---|---|---|
| ElevenLabs Sound Effects API β | Text β sound effect. "Ancient Roman marketplace at dusk" β audio | 10K free credits/mo (non-commercial) | Spatial awareness β understands "cathedral echo" vs "open field". Supports seamless_loop parameter |
| Freesound.org API | Search and stream CC-licensed ambient sounds | Free (API key required) | Massive library of real recordings |
| Pre-built ambient library | Bundle 10-15 ambient loops, select by visual_preset_key |
Free, zero latency | Instant, no API call needed |
| Web Audio API | Procedural ambient + reverb effects | Free, built into browser | ConvolverNode with impulse responses for environment-matched reverb (cathedral IR, cave IR, etc.) |
Recommended hybrid approach: Use ElevenLabs SFX API to generate destination-specific ambient audio on launch (it understands spatial context!), cache the results, and enhance with Web Audio API reverb/effects. Fall back to a pre-built ambient library for instant audio while the AI-generated soundscape loads.
Tier 2: Animated Talking Character (Medium Complexity)
This tier adds a character that moves and speaks β their lips sync to the audio, they have facial expressions.
2A. Real-Time Talking Avatar APIs (Cloud)
| Tool | What It Does | Pricing | Latency | Integration |
|---|---|---|---|---|
| Simli β | Image β animated talking head with lip sync via WebRTC. "Trinity" facial model with full expressions | $10 free credits + 50 free min/mo. Then ~$0.05/min | Sub-300ms | JS SDK, WebRTC. Best real-time option |
| D-ID Agents | Image β real-time avatar via WebRTC. Parallelized diffusion pipeline | 14-day trial (3 min, watermarked). From ~$4.70/mo | Sub-200ms, up to 100 FPS | Agents SDK with pre-built widget or full SDK |
| HeyGen LiveAvatar | Highest visual fidelity. "Avatar V" with fluid micro-expressions | Pay-as-you-go API credits (starts $5) | Low-latency streaming | WebRTC-based LiveAvatar API |
Recommended: Simli for real-time lip-synced avatars. It's (a) affordable with a real free tier, (b) truly real-time via WebRTC at sub-300ms, (c) takes a static image as input (feed it the AI-generated portrait from Tier 1), and (d) has a "Trinity" facial model with full expressions, not just basic lip movement. Your existing WebSocket architecture maps directly to this.
How Simli integrates with your architecture:
Your TTS audio β Simli WebRTC API β Lip-synced video frames β <video> element in browser
2B. Client-Side Lip Sync (FREE, No API Needed)
| Tool | What It Does | Real-Time? | Notes |
|---|---|---|---|
| TalkingHead (met4citizen) β | Complete JS class for 3D avatar rendering + lip-sync. Full-body GLB models with Mixamo animations. Built-in TTS integration (Google, Azure, ElevenLabs) | β Yes | Best complete free solution. One class handles loading, rendering, animating, and lip-syncing. Supports EN/DE/FR/FI/LT |
| Wawa-Lipsync β | Real-time viseme detection from any audio source, entirely in-browser. Uses Web Audio AnalyserNode β frequency β visemes |
β Yes | FREE npm package. Works with Three.js, R3F, or any framework. Lightweight and flexible |
| @beer-digital/lipsync-engine | Zero-dependency, renderer-agnostic viseme detection via AudioWorklet |
β Yes | Supports gapless streaming β important for streaming TTS output |
| Custom audio-reactive animation | Analyze TTS audio amplitude in real-time β animate portrait (pulse, glow, subtle scale, eye blink overlay) | β Yes | Zero cost, surprisingly effective. Works with your existing audio stream |
Rhubarb Lip Sync is NOT suitable for real-time β it's a CLI tool for offline/pre-recorded audio only. Don't use it for live conversation.
Quick win without full lip sync: Even animating the portrait based on audio amplitude creates a dramatically more alive feeling: pulse glow when speaking, subtle breathing idle animation, and programmatic eye blink overlay. This can be done in pure CSS/JS with your existing audio stream and is included in Tier 1's frontend overhaul.
Tier 3: Full 3D World (Highest Complexity, Maximum Wow)
3A. 360Β° Panoramic Scene Viewer
| Tool | What It Does | Notes |
|---|---|---|
| Blockade Labs + Three.js | Generate 360Β° equirectangular image, render as Three.js sphere interior. User can pan/look around | True "transported" feeling. Supports HDRI + 3D mesh export |
| A-Frame | Declarative 3D/VR framework. <a-sky src="scene.jpg"> β literally one line for a panorama |
Easiest 3D entry point |
| WebXR | VR headset support β user can enter the scene with Meta Quest, etc. | renderer.xr.enabled = true + VRButton in Three.js |
3B. 3D Character in Scene
| Tool | What It Does | Notes |
|---|---|---|
| MetaPerson (Avatar SDK) | Photo β full-body 3D avatar (GLB). Web, Unity, Unreal SDKs | Replaced Ready Player Me (which shut down Jan 31, 2026) |
| TalkingHead + MetaPerson | Generate 3D avatar β animate with TalkingHead class β lip sync with Wawa-Lipsync | Full pipeline, all free/open-source |
| Billboard in panorama | Place 2D character portrait as a billboard sprite in 3D scene, facing camera | Simpler than full 3D character, still immersive |
3C. Atmospheric 3D Effects
| Effect | Three.js Technique |
|---|---|
| Volumetric "God Rays" | Additive cone mesh + shadow maps |
| Height-based fog | TSL (Three.js Shading Language) custom fog |
| 50K+ particle systems | InstancedMesh + GPGPU compute |
| Bloom, chromatic aberration | RenderPipeline post-processing (r183+) |
| Atmospheric scattering | Rayleigh & Mie scattering shaders |
| Spatial audio | PannerNode with HRTF + ConvolverNode |
Ready Player Me shut down on January 31, 2026 (acquired by Netflix). All their SDKs are defunct. MetaPerson by Avatar SDK is the direct replacement for generating 3D avatars from photos.
Open Questions
Which tier do you want to start with? I recommend Tier 1 as it gives the biggest bang-for-buck and builds the foundation for Tier 2+.
Image generation API preference?
- Flux 2 via Together AI = zero new API keys, uses your existing account
- Blockade Labs = 360Β° panoramas for maximum immersion but new API signup + $20/mo
- Google Imagen 4 = high quality, $300 GCP trial credits
Budget sensitivity? Tier 1 is very cheap (pennies per image generation). Tier 2 (Simli) adds ~$0.05/min of conversation. Tier 3 is mostly free client-side tools.
Do you want to keep Gradio or are you open to a custom HTML frontend? Gradio constrains the UI significantly. Your current approach of injecting HTML/CSS/JS via
gr.HTML()works but is fighting Gradio's layout system. A standalone frontend (served by your existing FastAPI) would give total creative control. This is especially relevant for Tier 2+.Hackathon timeline? How much time do you have? Tier 1 could be done in 1-2 days. Tier 2 adds 1-2 more days. Tier 3 is 3+ days.
π° Cost Summary
| Service | Free Tier | Est. Monthly Cost |
|---|---|---|
| Flux 2 (via Together AI) | Same API key you already use | $5-20/month |
| Blockade Labs (360Β° scenes) | β | $20-48/month |
| ElevenLabs SFX | 10K credits/month (non-commercial) | $0-29/month |
| Freesound.org | Free | $0 |
| GSAP animations | Free forever | $0 |
| Three.js | Free forever | $0 |
| TalkingHead (3D avatar + lip sync) | Free forever | $0 |
| Wawa-Lipsync | Free forever | $0 |
| Simli (real-time avatar) | $10 credits + 50 min/mo | $0-50/month |
| D-ID Agents | 14-day trial only | $5-100/month |
| MetaPerson (3D avatar gen) | Free tier available | $0-20/month |
Proposed Changes (Tier 1 Implementation)
If you approve Tier 1, here's what I'd build:
Backend β Image Generation Service
[NEW] image_generation.py
New port interface SceneGenerator with methods: generate_scene(destination) β ImageResult and generate_portrait(persona, destination) β ImageResult
[NEW] image_gen_adapter.py
Implementation using Together AI Flux (or alternative). Builds rich prompts from destination/persona data.
[MODIFY] encounter_service.py
After generating destination and persona, call scene/portrait generation (can run in parallel). Emit new SceneImageEvent and PortraitImageEvent events.
[MODIFY] models.py
Add scene_image_url: str | None to EncounterSession (or a new VisualAssets model on the session).
[MODIFY] events.py
Add SceneImageEvent and PortraitImageEvent event types.
Frontend β Immersive Visual Layer
[MODIFY] cockpit.html
Add scene background container, character portrait <img>, particle <canvas>, ambient <audio> element.
[MODIFY] cockpit.css
Full-screen scene background styles, vignette overlay (radial-gradient), cinematic transitions, parallax effects, film grain (mix-blend-mode: overlay), era-specific color grading, floating conversation text, character portrait animations (breathing, glow pulse on speech).
[MODIFY] cockpit.js
Scene image loading and transitions, particle system engine, ambient audio management, character portrait audio-reactive animation, parallax mouse tracking, GSAP integration for launch sequence.
[NEW] ambience/
Bundled ambient audio loops (CC0): city, ocean, forest, desert, rain, fire, wind, marketplace, cathedral, etc.
Container & Config
[MODIFY] container.py
Wire up the new SceneGenerator adapter. Reuse existing TIME_MACHINE_LLM_API_KEY for Together AI image generation.
[MODIFY] view_models.py
Handle new image events, pass image URLs/base64 to frontend via Gradio state.
[MODIFY] realtime.py
Forward image events over WebSocket so the live voice UI also gets scene/portrait updates.
Verification Plan
Manual Verification
- Launch the app with
modalprofile - Click "Launch" with various coordinate prompts (e.g., "Ancient Rome", "1920s jazz club", "Mars colony 2847")
- Verify: scene background loads and fills screen with cinematic transition
- Verify: character portrait appears in aperture with breathing animation
- Verify: ambient audio plays matching the scene
- Verify: portrait pulses/glows when TTS audio plays
- Verify: conversation text appears as floating overlay near character
- Verify: film grain and color grading match the era
- Verify: particles match the environment (embers, snow, dust, etc.)
- Verify: works on both desktop and mobile viewports
- Verify:
prefers-reduced-motiondisables animations for accessibility
Automated Tests
- Unit test for prompt building (destination β image prompt)
- Unit test for ambient sound selection (destination keywords β sound file)
- Existing test suite still passes (
pytest)