sentinel / PLAYBOOK.md
kswffs's picture
Upload folder using huggingface_hub
b96103d verified
|
Raw
History Blame Contribute Delete
43 kB

A newer version of the Gradio SDK is available: 6.20.0

Upgrade

πŸ›‘οΈ Sentinel β€” Deep Implementation Playbook

10 Upgrades: Fully Architected, Battle-Planned

Rule: Agent does all the thinking. You review the plan, approve or redirect, then agent executes. Rule: Visual/UI work requires you to approve a reference style before agent writes any code.


UPGRADE 1: Client-Side Frame Differencing (OpenCV.js)

Goal: Browser analyzes frames locally, only uploads to server when motion detected. Files touched: sensor_bridge.py only. New lines: ~80 lines of JavaScript.

Architecture

CURRENT FLOW (wasteful):
  Camera β†’ canvas 320x240 β†’ base64 JPEG β†’ image-data textbox β†’ server process_frame
  Rate: 2 FPS always = 120 uploads/minute

NEW FLOW (smart):
  Camera β†’ canvas 320x240 β†’ OpenCV.js cv.absdiff(prev, current) β†’
  if changed_pixels > threshold β†’ upload base64 to server
  if not β†’ skip upload, increment local counter
  Rate: ~0.3-0.8 FPS average = 18-48 uploads/minute

Implementation Plan

Step 1: Add OpenCV.js CDN load to sensor_bridge.py

Insert at the top of the HTML string (before the IIFE):

<script async src="https://docs.opencv.org/4.x/opencv.js" 
        onload="window.__cvReady = true;" type="text/javascript"></script>

Step 2: Add state variables inside the IIFE

let prevMat = null;        // Previous frame as cv.Mat
let cvReady = false;       // OpenCV.js loaded flag
let skippedFrames = 0;     // Count of locally-filtered frames
const DIFF_THRESHOLD = 0.12; // 12% pixel change = significant

Step 3: Modify captureFrame() function

The current captureFrame() always uploads. New version:

captureFrame():
  1. Draw video to canvas (same as now)
  2. If cvReady:
     a. Read canvas into cv.Mat (srcMat)
     b. Convert to grayscale (cv.cvtColor)
     c. If prevMat exists:
        - cv.absdiff(prevMat, srcMat, diffMat)
        - cv.threshold(diffMat, threshMat, 25, 255, cv.THRESH_BINARY)
        - changedPct = cv.countNonZero(threshMat) / (width * height)
        - If changedPct < DIFF_THRESHOLD:
            skippedFrames++
            update UI indicator ("Filtered locally: X frames")
            cleanup Mats, return (DON'T upload)
     d. prevMat = srcMat.clone()
     e. cleanup temp Mats
  3. Proceed with base64 upload (same as now)

Step 4: Add skipped frame counter to UI

Add a hidden Gradio Number component:

# In app.py Monitor tab, alongside other hidden numbers:
skipped_frames = gr.Number(value=0, visible=False, elem_id="skipped-frames")

JS updates it via setGradioValue("skipped-frames", skippedFrames) every 2 seconds.

Edge Cases

  • OpenCV.js fails to load (CDN blocked): Fall back to current behavior (always upload). Check window.__cvReady before using cv.
  • First frame: No prevMat exists β†’ always upload first frame, store as prevMat.
  • Mat memory leak: Always call .delete() on temporary Mats. Only keep prevMat alive.
  • Camera off: Reset prevMat to null when camera stops.

Why This Wins

The 240x cost story goes from "trust our architecture" to "watch the counter β€” 847 frames processed locally, only 4 sent to GPU." That's a demo moment.

Testing Sequence

  1. Deploy with upgrade β†’ activate β†’ wave hand in front of camera β†’ see uploads spike
  2. Sit still β†’ see uploads drop to near zero
  3. Check that simulation mode still works (it bypasses camera entirely)

UPGRADE 2: Three.js / CSS3D Gyroscope HUD

Goal: Floating glass HUD overlay on camera feed. Tilts with phone. Shows live telemetry. Files touched: sensor_bridge.py (~200 lines JS/CSS). Dependency: None (pure client-side).

Architecture Decision

Option A β€” Pure CSS (Recommended for reliability):

  • No external library
  • CSS transform: perspective(1000px) rotateX() rotateY() driven by deviceorientation
  • ~80 lines, impossible to break
  • Looks great but not "3D"

Option B β€” Three.js CSS3DRenderer (Recommended for wow):

  • Load three.js from CDN (~150KB)
  • Create CSS3DObject from HUD HTML elements
  • Rotate entire CSS3D scene based on device orientation
  • ~200 lines, can break if CDN fails
  • True 3D parallax effect

My Recommendation: Option A for reliability, Option B for demo video

Start with Option A. It works on every device, including iOS. If you have time, add Option B as an enhancement.

Implementation Plan (Option A β€” Pure CSS)

Step 1: Add HUD container HTML to sensor_bridge.py

<div id="sentinel-hud" style="
    position: fixed;
    top: 80px;
    left: 50%;
    transform: translateX(-50%);
    z-index: 1000;
    pointer-events: none;
    transition: transform 0.1s ease-out;
    perspective: 800px;
    display: none;
">
    <div id="hud-inner" style="
        background: rgba(255,255,255,0.06);
        border: 1px solid rgba(255,255,255,0.15);
        border-radius: 16px;
        padding: 12px 20px;
        backdrop-filter: blur(20px);
        min-width: 240px;
        font-family: monospace;
        color: #a5b4fc;
    ">
        <div id="hud-compass">Heading: 0Β° N</div>
        <div id="hud-threat" style="color: #4ade80;">Threat: CLEAR</div>
        <div id="hud-fps">FPS: 2.0 | Filtered: 0</div>
        <div id="hud-battery">Battery: 100%</div>
    </div>
</div>

Step 2: Add deviceorientation handler

// Inside IIFE, after existing sensor handlers:
let hudActive = false;
const hudEl = document.getElementById("sentinel-hud");
const hudInner = document.getElementById("hud-inner");

function handleHUDOrientation(event) {
    if (!hudActive || !hudInner) return;
    // beta = front-back tilt (-180 to 180)
    // gamma = left-right tilt (-90 to 90)
    const tiltX = Math.max(-8, Math.min(8, event.gamma * 0.15));
    const tiltY = Math.max(-8, Math.min(8, (event.beta - 45) * 0.15));
    hudInner.style.transform = 
        `perspective(800px) rotateX(${-tiltY}deg) rotateY(${tiltX}deg)`;
}
window.addEventListener('deviceorientation', handleHUDOrientation);

Step 3: Update HUD data on each sensor tick

The existing sensor handlers already have access to heading, battery, etc. Add DOM updates:

// In the existing telemetry update cycle:
if (hudActive) {
    document.getElementById("hud-compass").textContent = 
        `Heading: ${heading}Β° ${getCompassDir(heading)}`;
    document.getElementById("hud-battery").textContent = 
        `Battery: ${batteryPct}%`;
    document.getElementById("hud-threat").textContent = 
        `Threat: ${currentThreatLevel}`;
    document.getElementById("hud-threat").style.color = 
        currentThreatLevel === "CLEAR" ? "#4ade80" : 
        currentThreatLevel === "WARNING" ? "#fbbf24" : "#ef4444";
}

Step 4: Show HUD on activation

In bindActivationButton(), after starting sensors:

if (hudEl) { hudEl.style.display = "block"; hudActive = true; }

What You Need to Approve

  • HUD position: top-center? bottom? full-width strip?
  • Data fields: compass + threat + FPS + battery? Or different data?
  • Glass style: subtle (0.06 opacity) or more visible (0.12)?
  • Color scheme: indigo text (#a5b4fc) or white?

Testing Sequence

  1. Activate on phone β†’ HUD appears floating over camera
  2. Tilt phone left β†’ HUD tilts right (parallax)
  3. Trigger simulation β†’ HUD threat level changes color
  4. Test on iOS β†’ works because it's pure CSS, no WebGL

UPGRADE 3: Binaural 3D Spatial Audio (Web Audio API)

Goal: Alert voice sounds like it's coming FROM the direction of danger. Files touched: sensor_bridge.py (100 lines JS), app.py (10 lines modified). Dependency: None (Web Audio API is native to all modern browsers).

Architecture

CURRENT:
  process_frame β†’ alert_html with <audio autoplay src="data:audio/wav;base64,...">
  β†’ browser plays audio from center (no spatial)

NEW:
  process_frame β†’ determines threat_x, threat_y from detection bbox center
  β†’ embeds them as data attributes in alert_html
  β†’ JS intercepts: decodes base64 audio β†’ creates AudioBufferSourceNode
  β†’ connects to PannerNode (HRTF model, positioned at threat_x/y/z)
  β†’ connects to AudioListener (oriented by deviceorientation)
  β†’ audio plays from spatial position

Implementation Plan

Step 1: Add spatial audio engine to sensor_bridge.py

// Inside IIFE:
let spatialAudioCtx = null;
let audioListener = null;

function initSpatialAudio() {
    spatialAudioCtx = new (window.AudioContext || window.webkitAudioContext)();
    audioListener = spatialAudioCtx.listener;
    // Set listener to face forward
    if (audioListener.positionX) {
        audioListener.positionX.value = 0;
        audioListener.positionY.value = 0;
        audioListener.positionZ.value = 0;
        audioListener.forwardX.value = 0;
        audioListener.forwardY.value = 0;
        audioListener.forwardZ.value = -1;
    }
}

async function playSpatialAudio(base64Audio, threatX, threatY) {
    if (!spatialAudioCtx) initSpatialAudio();
    if (spatialAudioCtx.state === 'suspended') await spatialAudioCtx.resume();

    // Decode base64 WAV to AudioBuffer
    const binary = atob(base64Audio);
    const bytes = new Uint8Array(binary.length);
    for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i);
    const audioBuffer = await spatialAudioCtx.decodeAudioData(bytes.buffer);

    // Create source
    const source = spatialAudioCtx.createBufferSource();
    source.buffer = audioBuffer;

    // Create 3D panner
    const panner = spatialAudioCtx.createPanner();
    panner.panningModel = 'HRTF';
    panner.distanceModel = 'inverse';
    panner.refDistance = 1;
    panner.maxDistance = 10;
    // Position: threatX maps to left(-1) to right(+1), Y to up/down
    panner.positionX.value = threatX * 2;  // -2 to +2 meters
    panner.positionY.value = threatY;
    panner.positionZ.value = -2;  // 2 meters in front

    // Connect: source β†’ panner β†’ destination
    source.connect(panner);
    panner.connect(spatialAudioCtx.destination);
    source.start(0);
}

Step 2: Update deviceorientation to rotate listener

// In existing handleOrientation function, add:
if (spatialAudioCtx && audioListener) {
    const rad = (event.alpha || 0) * Math.PI / 180;
    if (audioListener.forwardX) {
        audioListener.forwardX.value = -Math.sin(rad);
        audioListener.forwardZ.value = -Math.cos(rad);
    }
}

Step 3: Modify app.py alert generation

In process_frame, when building alert_html, embed threat coordinates:

# After YOLO detections, calculate threat position
threat_x = 0.0  # center by default
threat_y = 0.0
if detections:
    # Use center of first trigger detection bbox, normalized to -1..1
    det = detections[0]
    threat_x = ((det.bbox[0] + det.bbox[2]) / 2 / 640 - 0.5) * 2  # -1 to 1
    threat_y = -((det.bbox[1] + det.bbox[3]) / 2 / 480 - 0.5) * 2  # inverted Y

# In alert_html, add data attributes:
alert_html = f"""<div class='alert-banner-{alert_level}' 
    data-threat-x='{threat_x:.2f}' 
    data-threat-y='{threat_y:.2f}'
    data-audio-src='{audio_data_uri}'>
    ⚠️ {alert_text}{alert_image_html}</div>"""
# Remove the <audio autoplay> tag β€” JS handles playback now

Step 4: JS observer detects new alert and triggers spatial playback

// MutationObserver on alert-banner-container
const alertObserver = new MutationObserver((mutations) => {
    for (const m of mutations) {
        const banner = m.target.querySelector('[data-audio-src]');
        if (banner) {
            const audioSrc = banner.dataset.audioSrc;
            const tx = parseFloat(banner.dataset.threatX || 0);
            const ty = parseFloat(banner.dataset.threatY || 0);
            if (audioSrc && audioSrc.startsWith('data:audio')) {
                const base64 = audioSrc.split(',')[1];
                playSpatialAudio(base64, tx, ty);
            }
        }
    }
});

What You Need to Approve

  • Fallback: if Web Audio fails (old browser), fall back to <audio autoplay> tag?
  • Spatial range: Β±2 meters (subtle) or Β±5 meters (dramatic)?
  • Should listener rotate with gyroscope or stay fixed forward?

Testing Sequence

  1. Put on headphones
  2. Trigger "Person Approaching" sim (person comes from left)
  3. Voice should sound like it's coming from the LEFT earphone
  4. Turn phone 90Β° β†’ voice position shifts relative to your head
  5. Test without headphones β†’ still works but less immersive

UPGRADE 4: Dynamic SVG Bounding Box Overlays

Goal: YOLO detection boxes drawn as animated SVG overlays on the camera feed. Files touched: sensor_bridge.py (50 lines JS), app.py (30 lines). Dependency: Upgrade 1 (client-side frame diff) recommended first.

Architecture

CURRENT:
  process_frame β†’ YOLO detect β†’ detections used internally only β†’ discarded
  Camera feed shows raw video with no annotations

NEW:
  process_frame β†’ YOLO detect β†’ serialize detections as JSON β†’
  write to hidden textbox "detection-data" β†’
  JS reads JSON β†’ creates/updates SVG overlay on camera feed β†’
  animated rectangles with class labels

Implementation Plan

Step 1: Always run YOLO (not just when gatekeeper triggers)

In app.py, change:

# BEFORE (line 563):
if change_result.is_significant and object_detector is not None:

# AFTER:
if object_detector is not None:

This means YOLO runs on every frame, not just significant changes. Cost: ~5ms per frame on CPU. Acceptable for real-time display. The GATEKEEPER still decides whether to call VLM β€” this only changes what we SHOW.

Step 2: Serialize detections to JSON

After YOLO runs, write detection data:

# After line 565 (yolo_summary = ...):
detection_json = json.dumps([
    {
        "class": d.class_name,
        "conf": round(d.confidence, 2),
        "bbox": list(d.bbox),
        "trigger": d.is_trigger
    }
    for d in detections
])

Step 3: Add hidden textbox for detection data

In app.py Monitor tab:

detection_data = gr.Textbox(value="[]", visible=False, elem_id="detection-data")

Update it in process_frame return tuple (add to outputs).

Step 4: SVG overlay in sensor_bridge.py

// Create SVG overlay positioned over the camera feed image
function createSVGOverlay() {
    const cameraImg = document.querySelector('#alert-banner-container')
        ?.closest('.glass-panel')
        ?.querySelector('img');  // The camera feed <img> element
    if (!cameraImg) return null;
    
    const svg = document.createElementNS("http://www.w3.org/2000/svg", "svg");
    svg.id = "detection-svg";
    svg.style.cssText = "position:absolute;top:0;left:0;width:100%;height:100%;pointer-events:none;z-index:10;";
    cameraImg.parentElement.style.position = "relative";
    cameraImg.parentElement.appendChild(svg);
    return svg;
}

// Color map for detection classes
const CLASS_COLORS = {
    "person": "#3b82f6",
    "vehicle": "#ef4444", 
    "fire": "#f97316",
    "animal": "#a855f7",
    "default": "#22c55e"
};

function drawDetections(detections) {
    const svg = document.getElementById("detection-svg") || createSVGOverlay();
    if (!svg) return;
    svg.innerHTML = ""; // Clear previous

    detections.forEach(det => {
        const [x1, y1, x2, y2] = det.bbox;
        const color = CLASS_COLORS[det.class] || CLASS_COLORS.default;
        const w = x2 - x1, h = y2 - y1;

        // Animated rectangle
        const rect = document.createElementNS("http://www.w3.org/2000/svg", "rect");
        rect.setAttribute("x", x1); rect.setAttribute("y", y1);
        rect.setAttribute("width", w); rect.setAttribute("height", h);
        rect.setAttribute("fill", "none");
        rect.setAttribute("stroke", color);
        rect.setAttribute("stroke-width", "3");
        rect.setAttribute("rx", "4");
        rect.style.opacity = "0";
        rect.style.transition = "opacity 0.3s";
        svg.appendChild(rect);
        requestAnimationFrame(() => rect.style.opacity = "0.8");

        // Label
        const text = document.createElementNS("http://www.w3.org/2000/svg", "text");
        text.setAttribute("x", x1); text.setAttribute("y", y1 - 6);
        text.setAttribute("fill", color);
        text.setAttribute("font-size", "14");
        text.setAttribute("font-family", "monospace");
        text.textContent = `${det.class} ${(det.conf * 100).toFixed(0)}%`;
        svg.appendChild(text);
    });
}

Step 5: Observe detection-data changes

// MutationObserver on hidden detection-data textbox
const detObserver = new MutationObserver(() => {
    const el = document.querySelector('#detection-data input, #detection-data textarea');
    if (!el) return;
    try {
        const dets = JSON.parse(el.value || "[]");
        drawDetections(dets);
    } catch(e) {}
});

What You Need to Approve

  • Box style: full rectangle? corner brackets? with fill?
  • Corner brackets look more "professional AI" (think Tesla vision display)
  • Colors: class-based (above) or all same color?
  • Always show boxes or only on trigger detections?

Testing Sequence

  1. Activate β†’ point camera at a person β†’ blue box appears with "person 87%"
  2. Point at a car β†’ red box "vehicle 92%"
  3. Empty room β†’ no boxes
  4. Verify boxes scale correctly when window is resized

UPGRADE 5: Kokoro TTS Streaming

Goal: First word of alert audio plays within 300ms instead of waiting 2+ seconds for full generation. Files touched: kokoro_tts.py (40 lines), app.py (20 lines). Dependency: None.

Architecture

CURRENT:
  speaker.speak(full_text) β†’ Kokoro generates complete WAV β†’ base64 encode β†’ return
  Latency: 1.5-3.0 seconds for a 2-sentence alert

NEW:
  Split text into sentences β†’ generate sentence 1 audio β†’ yield immediately
  β†’ generate sentence 2 audio β†’ yield
  Client receives sentence 1 audio at ~300ms, starts playing
  Sentence 2 arrives while sentence 1 is still playing

Implementation Plan

Step 1: Add streaming method to AlertSpeaker

In kokoro_tts.py, add a new method:

async def speak_streaming(self, text: str, level: str = "warning"):
    """
    Generator that yields audio chunks sentence by sentence.
    Each chunk is a base64-encoded WAV.
    """
    import re
    sentences = re.split(r'(?<=[.!?])\s+', text.strip())
    
    for sentence in sentences:
        if not sentence.strip():
            continue
        try:
            result = await self.speak(sentence, level=level)
            if result.get("audio_base64"):
                yield result["audio_base64"]
        except Exception as e:
            logger.error("TTS streaming chunk failed", error=str(e))
            break

Step 2: Modify app.py to use streaming

Instead of generating full audio then embedding in HTML, use Gradio's streaming audio:

# Replace the current TTS section in process_frame with:
audio_chunks = []
if alert_level in ["critical", "warning"]:
    async for chunk_b64 in speaker.speak_streaming(alert_text, level=alert_level):
        audio_chunks.append(chunk_b64)
    
    # Use first chunk for immediate playback, concatenate for archive
    if audio_chunks:
        audio_data_uri = f"data:audio/wav;base64,{audio_chunks[0]}"
        # Pass remaining chunks via hidden state for sequential playback

Step 3: Add sequential audio playback in JS

In sensor_bridge.py:

// Queue-based audio player that chains chunks
let audioQueue = [];
let isPlaying = false;

function queueAudio(base64Chunks) {
    audioQueue = [...base64Chunks];
    if (!isPlaying) playNext();
}

async function playNext() {
    if (audioQueue.length === 0) { isPlaying = false; return; }
    isPlaying = true;
    const chunk = audioQueue.shift();
    const audio = new Audio(`data:audio/wav;base64,${chunk}`);
    audio.onended = () => playNext();
    await audio.play();
}

Edge Cases

  • Single sentence alerts: Works normally, one chunk, no difference from current.
  • Kokoro fails on sentence 2: Graceful stop β€” sentence 1 already played.
  • User triggers new alert while audio playing: Clear queue, start new alert audio.

Testing Sequence

  1. Trigger "Fall Detected" scenario (multi-sentence alert)
  2. Note: first word plays within ~300ms
  3. Second sentence plays seamlessly after first
  4. Trigger another alert mid-playback β†’ old audio stops, new starts

UPGRADE 6: IMU-Linked Glassmorphic UI

Goal: Glass panels shift highlights and shadows based on phone tilt. Feels like real glass. Files touched: sensor_bridge.py (40 lines JS), app.py CUSTOM_CSS (15 lines). Dependency: None. This is the simplest visual upgrade.

Architecture

deviceorientation event (60 FPS) β†’
  CSS custom properties on document.documentElement:
    --tilt-x: gamma * 0.3  (left-right)
    --tilt-y: beta * 0.3   (front-back)
  
CSS uses these properties:
  .glass-panel::before {
    background: radial-gradient(
      circle at calc(50% + var(--tilt-x) * 1px) calc(50% + var(--tilt-y) * 1px),
      rgba(255,255,255,0.08) 0%,
      transparent 60%
    );
  }
  .glass-panel {
    box-shadow: 
      calc(var(--tilt-x) * -0.5px) calc(var(--tilt-y) * -0.5px) 20px rgba(0,0,0,0.3);
  }

Implementation Plan

Step 1: Add CSS custom properties to CUSTOM_CSS in app.py

:root {
    --tilt-x: 0;
    --tilt-y: 0;
}
.glass-panel {
    position: relative;
    overflow: hidden;
    transition: box-shadow 0.15s ease-out;
    box-shadow: 
        calc(var(--tilt-x) * -0.5px) 
        calc(var(--tilt-y) * -0.5px) 
        24px rgba(0,0,0,0.35);
}
.glass-panel::before {
    content: '';
    position: absolute;
    top: 0; left: 0; right: 0; bottom: 0;
    background: radial-gradient(
        circle at calc(50% + var(--tilt-x) * 2px) calc(30% + var(--tilt-y) * 2px),
        rgba(255,255,255,0.07) 0%,
        transparent 50%
    );
    pointer-events: none;
    z-index: 1;
    transition: background 0.15s ease-out;
}

Step 2: Add JS tilt handler to sensor_bridge.py

// Inside IIFE, add alongside existing deviceorientation handler:
const MAX_TILT = 25; // cap at Β±25px shift
function handleGlassTilt(event) {
    const gamma = Math.max(-30, Math.min(30, event.gamma || 0));
    const beta = Math.max(-30, Math.min(30, (event.beta || 45) - 45));
    document.documentElement.style.setProperty('--tilt-x', gamma * 0.8);
    document.documentElement.style.setProperty('--tilt-y', beta * 0.8);
}
window.addEventListener('deviceorientation', handleGlassTilt);

What You Need to Approve

  • Tilt sensitivity: 0.8 multiplier (moderate)? 0.3 (subtle)? 1.5 (dramatic)?
  • Highlight intensity: 0.07 opacity (shown above)? Higher = more visible but less classy
  • Desktop fallback: track mouse position instead of gyroscope? Or no effect on desktop?

Testing Sequence

  1. Open on phone β†’ panels look normal
  2. Tilt phone left β†’ highlight slides right, shadow shifts left
  3. Tilt forward β†’ highlight moves up, shadow moves down
  4. Looks like a physical piece of glass catching light
  5. On desktop with no gyro β†’ either static or mouse-follows

UPGRADE 7: Live Cost Race Ticker

Goal: Split-screen widget showing Sentinel vs GPT-4o costs racing in real-time. Files touched: sensor_bridge.py (100 lines JS/HTML), app.py (20 lines). Dependency: None.

Architecture

JS reads cost_tracker data from hidden Gradio components:
  - cost-total (existing) β†’ Sentinel's actual cost
  - frames-processed (calculate from frame_count_state)

Calculates naive GPT-4o cost:
  frames_processed * $0.01 per frame (GPT-4o Vision at 2 FPS)

Displays two animated counters:
  LEFT:  Sentinel $0.0023 (green, slow counter)
  RIGHT: GPT-4o  $1.4200 (red, fast counter)
  BOTTOM: "You saved $1.42 (99.8%)"

Implementation Plan

Step 1: Add cost ticker HTML to Monitor tab

In sensor_bridge.py, add HTML that gets injected:

<div id="cost-ticker" style="
    display: none;
    margin-top: 12px;
    padding: 14px;
    background: rgba(0,0,0,0.3);
    border-radius: 12px;
    border: 1px solid rgba(255,255,255,0.08);
    font-family: monospace;
">
    <div style="display: flex; justify-content: space-between; align-items: center;">
        <div style="text-align: center;">
            <div style="color: #4ade80; font-size: 11px;">SENTINEL</div>
            <div id="ticker-sentinel" style="color: #4ade80; font-size: 22px; font-weight: bold;">$0.0000</div>
        </div>
        <div style="color: #6b7280; font-size: 20px;">vs</div>
        <div style="text-align: center;">
            <div style="color: #ef4444; font-size: 11px;">GPT-4o VISION</div>
            <div id="ticker-gpt4o" style="color: #ef4444; font-size: 22px; font-weight: bold;">$0.0000</div>
        </div>
    </div>
    <div id="ticker-savings" style="
        text-align: center;
        margin-top: 8px;
        color: #a5b4fc;
        font-size: 13px;
    ">Savings: $0.00 (0%)</div>
    <div style="
        height: 4px;
        background: rgba(255,255,255,0.1);
        border-radius: 2px;
        margin-top: 8px;
        overflow: hidden;
    ">
        <div id="ticker-bar" style="
            height: 100%;
            background: linear-gradient(90deg, #4ade80, #22d3ee);
            border-radius: 2px;
            width: 0%;
            transition: width 0.5s ease-out;
        "></div>
    </div>
</div>

Step 2: Add animation JS

// Inside IIFE:
const GPT4O_COST_PER_FRAME = 0.01; // $0.01 per frame at 2 FPS
let displayedSentinel = 0;
let displayedGPT4o = 0;

function updateCostTicker() {
    if (!tickerActive) return;
    
    // Read actual values from Gradio state
    const sentinelCost = parseFloat(
        document.querySelector('#cost-total input')?.value || '0'
    );
    const framesProcessed = parseInt(
        document.querySelector('#frame-count input')?.value || '0'
    );
    const gpt4oCost = framesProcessed * GPT4O_COST_PER_FRAME;
    
    // Animate numbers toward target (easing)
    displayedSentinel += (sentinelCost - displayedSentinel) * 0.15;
    displayedGPT4o += (gpt4oCost - displayedGPT4o) * 0.15;
    
    document.getElementById("ticker-sentinel").textContent = 
        `$${displayedSentinel.toFixed(4)}`;
    document.getElementById("ticker-gpt4o").textContent = 
        `$${displayedGPT4o.toFixed(2)}`;
    
    const savings = displayedGPT4o - displayedSentinel;
    const savingsPct = displayedGPT4o > 0 ? (savings / displayedGPT4o * 100) : 0;
    document.getElementById("ticker-savings").textContent = 
        `You saved $${savings.toFixed(2)} (${savingsPct.toFixed(1)}%)`;
    
    // Update progress bar
    const barWidth = Math.min(100, savingsPct);
    document.getElementById("ticker-bar").style.width = `${barWidth}%`;
}

// Run at 30 FPS for smooth animation
setInterval(updateCostTicker, 33);

Step 3: Expose frame count to JS

Add hidden component in app.py:

frame_count_display = gr.Number(value=0, visible=False, elem_id="frame-count")

Update it in process_frame return or via the existing frame_count_state.

Step 4: Show ticker on activation

In activation handler: document.getElementById("cost-ticker").style.display = "block";

What You Need to Approve

  • GPT-4o cost rate: $0.01/frame is approximate. Use different number?
  • Number animation: smooth easing (shown) or instant snap?
  • Should there be confetti when savings hits $1.00?
  • Placement: inside glass-panel below camera? or separate section?

Testing Sequence

  1. Activate β†’ ticker shows $0.00 vs $0.00
  2. Run 5 simulation scenarios β†’ Sentinel: ~$0.0015 vs GPT-4o: ~$0.05
  3. Watch numbers animate smoothly
  4. Progress bar fills as savings grow
  5. After 50 frames: "$0.49 saved (99.8%)" β€” powerful demo moment

UPGRADE 8: Multi-Turn VLM Memory

Goal: Nemotron sees last 3 alerts and reasons about trends. "Person was 5m away, now 1m." Files touched: app.py only (~30 lines). Dependency: None.

Architecture

NEW gr.State: vlm_memory = gr.State(value=[])

On each VLM call:
  1. Append current alert context to vlm_memory (max 3 entries)
  2. Format memory into system prompt:
     "Previous alerts in this session:
      [10:05:01] WARNING: Person detected 5m away
      [10:05:11] WARNING: Person now 3m away, approaching
      [current] Analyze this new frame..."
  3. Nemotron receives history + current frame β†’ can reference trends
  4. Trim memory to last 3 entries (sliding window)

Implementation Plan

Step 1: Add memory state

In app.py, with other gr.State declarations:

vlm_memory = gr.State(value=[])  # List of dicts: {time, level, text}

Step 2: Update SYSTEM_PROMPT

Change from static string to a function:

def build_system_prompt(memory: list) -> str:
    history_block = ""
    if memory:
        lines = [f"  [{m['time']}] {m['level']}: {m['text']}" for m in memory[-3:]]
        history_block = "\n\nPrevious alerts in this session (use for trend awareness):\n" + "\n".join(lines)
    
    return f"""You are Sentinel, an autonomous AI guardian for visually impaired and elderly users.
You receive visual descriptions and sensor data. Your job is to:
1. Identify potential dangers (tripping hazards, approaching vehicles, strangers, fire)
2. Provide navigation guidance (door ahead, stairs, obstacles)
3. Alert ONLY when genuinely dangerous β€” avoid false alarms
4. Respond in 1-2 sentences maximum (user hears this via TTS)
5. Reference previous alerts when relevant (e.g., "Person getting closer")
6. Escalate if a threat is worsening over time

Format: [LEVEL] message
Where LEVEL is: CRITICAL, WARNING, or OK{history_block}"""

Step 3: Update process_frame

After receiving VLM response and parsing alert:

# Append to memory
if alert_level in ["critical", "warning"]:
    memory_entry = {
        "time": time.strftime("%H:%M:%S"),
        "level": alert_level.upper(),
        "text": alert_text[:80]  # Truncate to save tokens
    }
    vlm_memory.append(memory_entry)
    vlm_memory = vlm_memory[-3:]  # Keep last 3 only

# Use dynamic prompt:
question = f"Analyze this scene. Context: {sensor_context}"
# Pass build_system_prompt(vlm_memory) instead of static SYSTEM_PROMPT

Step 4: Update function signature and event bindings

Add vlm_memory to process_frame params (now 21 params, 15 returns, 9 gr.State). Update all return paths and image_data.change bindings accordingly.

Token Budget

  • Each memory entry: ~20 tokens
  • 3 entries: ~60 extra tokens per VLM call
  • Negligible cost increase, significant intelligence increase

Testing Sequence

  1. Trigger "Person Approaching" 3 times in sequence
  2. First alert: "Person detected 5m away"
  3. Second alert: "Person closer, now 3m" (should reference previous)
  4. Third alert: "CRITICAL: Person very close, 1m away β€” was 5m moments ago"
  5. Verify memory shows in alert history with trend references

UPGRADE 9: Virtual Walk Mode

Goal: Pre-recorded walking video plays. Sentinel analyzes it live. Alerts interrupt the video. Files touched: app.py (80 lines new tab), sensor_bridge.py (60 lines JS). Dependency: Upgrade 3 (spatial audio) recommended for full experience.

Architecture

New Tab: "🎬 Demo"

Components:
  - gr.Video (pre-loaded walking video)
  - gr.HTML (alert overlay zone, positioned over video)
  - Play/Pause button
  - Speed control (0.5x, 1x, 2x)

JS Frame Extractor:
  - requestVideoFrameCallback on video element
  - Every 1000ms (1 FPS): draw video frame to canvas
  - Convert to base64
  - Inject into image-data textbox β†’ triggers process_frame pipeline
  
Alert Overlay:
  - When alert generated, HTML overlay appears ON TOP of video
  - Semi-transparent banner with alert text
  - Auto-dismisses after 5 seconds
  - Audio plays via spatial audio (Upgrade 3) or regular playback

Data Flow:
  video.play() β†’ JS extracts frame every 1s β†’ base64 β†’ image-data
  β†’ process_frame β†’ VLM β†’ alert β†’ overlay on video + TTS
  β†’ video continues playing underneath

Implementation Plan

Step 1: Create Demo tab in app.py

with gr.Tab("🎬 Demo"):
    with gr.Column(elem_classes="glass-panel"):
        gr.Markdown("### Virtual Walk Mode")
        gr.Markdown("Watch Sentinel analyze a real walking scenario in real-time.")
        
        demo_video = gr.Video(
            label="Walking Scenario",
            value=None,  # User uploads or pre-load
            autoplay=False,
            elem_id="demo-video"
        )
        
        # Pre-loaded scenario buttons
        with gr.Row():
            gr.Markdown("**Or try a preset scenario:**")
        with gr.Row():
            demo_stairs_btn = gr.Button("Stairs", size="sm")
            demo_traffic_btn = gr.Button("Traffic", size="sm")
            demo_park_btn = gr.Button("Park Walk", size="sm")
        
        demo_alert_overlay = gr.HTML(
            "",
            elem_id="demo-alert-overlay"
        )
        
        demo_status = gr.HTML(
            "<div style='text-align:center;color:#9ca3af;'>Press Play to start the virtual walk</div>",
            elem_id="demo-status"
        )

Step 2: JS frame extractor in sensor_bridge.py

function initDemoMode() {
    const video = document.querySelector('#demo-video video');
    if (!video) return;
    
    let demoCanvas = document.createElement('canvas');
    demoCanvas.width = 640;
    demoCanvas.height = 480;
    let demoCtx = demoCanvas.getContext('2d');
    let lastDemoFrame = 0;
    
    video.addEventListener('play', () => {
        // Activate monitoring for demo mode
        setGradioValue("sentinel-active-state", "true");
    });
    
    function extractDemoFrame(timestamp) {
        if (video.paused || video.ended) return;
        
        if (timestamp - lastDemoFrame >= 1000) { // 1 FPS
            lastDemoFrame = timestamp;
            demoCtx.drawImage(video, 0, 0, 640, 480);
            const base64 = demoCanvas.toDataURL("image/jpeg", 0.7).split(",")[1];
            setGradioValue("image-data", base64);
        }
        requestAnimationFrame(extractDemoFrame);
    }
    
    video.addEventListener('play', () => {
        requestAnimationFrame(extractDemoFrame);
    });
}

Step 3: Alert overlay CSS

#demo-alert-overlay > div {
    position: absolute;
    bottom: 80px;
    left: 50%;
    transform: translateX(-50%);
    z-index: 100;
    max-width: 90%;
    animation: slideUp 0.3s ease-out;
}
@keyframes slideUp {
    from { opacity: 0; transform: translateX(-50%) translateY(20px); }
    to { opacity: 1; transform: translateX(-50%) translateY(0); }
}

What You Need to Decide

  • Video source: Film your own walking video? Use Creative Commons? Use a slideshow of sim images instead?
  • Frame rate: 1 FPS (1 call/sec, cheap) or 2 FPS?
  • Should alerts PAUSE the video or overlay while playing?
  • Preset videos: record 3 short clips (stairs, traffic, park) and embed?

Testing Sequence

  1. Upload walking video to Demo tab
  2. Press Play β†’ video starts, frame counter ticks
  3. When person crosses path β†’ alert banner slides up over video + TTS plays
  4. When stairs appear β†’ CRITICAL alert + loud warning
  5. Clear section β†’ "OK: Path clear" in info banner

UPGRADE 10: Web Speech API Panic Override

Goal: Browser listens for "HELP"/"STOP" keywords. Works offline. Triggers emergency mode. Files touched: sensor_bridge.py (~50 lines JS). Dependency: None.

Architecture

webkitSpeechRecognition runs continuously in browser:
  - Listens for keywords: "help", "stop", "emergency", "cancel"
  - On match:
    1. Set monitoring_active to true (via hidden component)
    2. Inject CRITICAL alert into alert banner
    3. Play emergency siren sound (Web Audio oscillator)
    4. Open SMS app with pre-filled GPS + "EMERGENCY" text
  - Works even if server is unreachable
  - iOS requires user gesture to start (use activation button)

Implementation Plan

Step 1: Add speech recognition to sensor_bridge.py IIFE

// Inside IIFE:
let recognition = null;
const PANIC_KEYWORDS = ["help", "stop", "emergency", "cancel", "sos"];

function initSpeechPanic() {
    const SpeechRecognition = window.SpeechRecognition || window.webkitSpeechRecognition;
    if (!SpeechRecognition) {
        console.log("Web Speech API not available. Panic override disabled.");
        return;
    }
    
    recognition = new SpeechRecognition();
    recognition.continuous = true;
    recognition.interimResults = true;
    recognition.lang = 'en-US';
    
    recognition.onresult = (event) => {
        const last = event.results[event.results.length - 1];
        const transcript = last[0].transcript.toLowerCase().trim();
        
        for (const keyword of PANIC_KEYWORDS) {
            if (transcript.includes(keyword)) {
                triggerPanicMode(keyword);
                break;
            }
        }
    };
    
    recognition.onerror = (event) => {
        // Auto-restart on error (except if user stopped it)
        if (event.error !== 'aborted' && isActivated) {
            setTimeout(() => { try { recognition.start(); } catch(e) {} }, 1000);
        }
    };
    
    recognition.onend = () => {
        // Auto-restart if still active
        if (isActivated) {
            try { recognition.start(); } catch(e) {}
        }
    };
}

function triggerPanicMode(keyword) {
    console.log(`PANIC KEYWORD DETECTED: "${keyword}"`);
    
    // 1. Set alert banner to critical
    const banner = document.getElementById("alert-banner-container");
    if (banner) {
        banner.innerHTML = `
            <div class="alert-banner-critical" style="animation: pulse 0.5s infinite;">
                🚨 EMERGENCY MODE ACTIVATED<br>
                Voice command: "${keyword.toUpperCase()}"<br>
                <small>GPS coordinates being sent to emergency contact</small>
            </div>`;
    }
    
    // 2. Play emergency siren (Web Audio oscillator)
    playEmergencySiren();
    
    // 3. Open SMS with pre-filled message
    const lat = document.querySelector('#gps-lat input')?.value || 'unknown';
    const lon = document.querySelector('#gps-lon input')?.value || 'unknown';
    const smsBody = encodeURIComponent(
        `SENTINEL EMERGENCY: User triggered panic mode. Last known location: ${lat}, ${lon}. Please check on them.`
    );
    window.open(`sms:?body=${smsBody}`, '_blank');
}

function playEmergencySiren() {
    const ctx = new (window.AudioContext || window.webkitAudioContext)();
    const osc = ctx.createOscillator();
    const gain = ctx.createGain();
    osc.type = 'sawtooth';
    osc.frequency.value = 800;
    gain.gain.value = 0.3;
    osc.connect(gain);
    gain.connect(ctx.destination);
    osc.start();
    // Modulate frequency for siren effect
    osc.frequency.setValueAtTime(800, ctx.currentTime);
    osc.frequency.linearRampToValueAtTime(1200, ctx.currentTime + 0.5);
    osc.frequency.linearRampToValueAtTime(800, ctx.currentTime + 1.0);
    // Auto-stop after 3 seconds
    gain.gain.linearRampToValueAtTime(0, ctx.currentTime + 3.0);
    osc.stop(ctx.currentTime + 3.0);
}

Step 2: Start recognition on activation

In startAllSensors(), add: initSpeechPanic(); if (recognition) try { recognition.start(); } catch(e) {}

In stopAllSensors(), add: if (recognition) try { recognition.stop(); } catch(e) {}

Edge Cases

  • iOS Safari: Requires user gesture. Since activation button is the user gesture, speech recognition starts after clicking ACTIVATE.
  • Background noise false triggers: Use interimResults: true but only trigger on high-confidence matches. Could add confidence threshold: if (last[0].confidence > 0.7).
  • Browser doesn't support: Graceful log message, no crash.

Testing Sequence

  1. Activate Sentinel
  2. Say "Help" clearly β†’ banner turns red, siren plays, SMS app opens
  3. Say "Stop" β†’ same emergency response
  4. Normal conversation β†’ no false triggers
  5. Test with server offline β†’ still works (pure browser-side)

BUILD ORDER & TIMELINE

Day 1 β€” Deploy + Visual Wins (6 hours)

Time Upgrade Why First
1h Deploy Modal + HF Space Everything depends on this
1h Upgrade 6: IMU Glass UI Simplest visual upgrade, 40 lines
1h Upgrade 7: Cost Race Ticker Most persuasive demo element
1h Upgrade 4: SVG Bounding Boxes Camera feed looks like real AI
1h Upgrade 2: Gyroscope HUD Premium app feel
1h Quick wins (tab badges, keyboard shortcuts) Polish

Day 2 β€” Intelligence + Audio (5 hours)

Time Upgrade Why This Order
1h Upgrade 8: Multi-Turn VLM Memory Pure app.py change, smart
1h Upgrade 1: OpenCV.js Frame Diff Architecture proof
1.5h Upgrade 3: Spatial Audio Demo-defining feature
1h Upgrade 5: TTS Streaming Pairs with spatial audio
0.5h Upgrade 10: Speech Panic Override Simple, impressive

Day 3 β€” Killer Demo + Submission (5 hours)

Time Task Notes
1h Upgrade 9: Virtual Walk Mode Record walking video first
1h Record demo video Show all upgrades working together
1h HF Dataset (500 rows) Export from simulation + demo runs
1h Space README + thumbnail Architecture diagram, screenshots
1h Submit + blog post Link everything

DEPENDENCY MAP

No dependencies:
  Upgrade 2 (HUD) ──────────────────────┐
  Upgrade 6 (Glass UI) ─────────────────
  Upgrade 7 (Cost Ticker) ──────────────
  Upgrade 8 (VLM Memory) ──────────────
  Upgrade 10 (Speech Panic) ────────────
                                        β”‚
Optional dependencies:                  β”‚
  Upgrade 1 (Frame Diff) β†’ Upgrade 4 (SVG Boxes)
  Upgrade 3 (Spatial Audio) β†’ Upgrade 9 (Virtual Walk)
  Upgrade 5 (TTS Stream) β†’ Upgrade 3 (Spatial Audio)
                                        β”‚
All feed into: β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  Upgrade 9 (Virtual Walk) ← uses everything

WHAT YOU APPROVE BEFORE AGENT BUILDS

For each upgrade, tell the agent:

  1. "Build Upgrade X" β€” agent starts
  2. "Use approach Y" β€” if there are options (like HUD Option A vs B)
  3. "Reference: [screenshot/description]" β€” for visual upgrades, describe what you want
  4. "Skip the [thing I don't want]" β€” agent trims unnecessary parts

The agent has all the architecture. It just needs your green light on style choices.