agent-tina / docs /progress.md
KPrashanth's picture
Deploy Agent Tina
ab74a45 verified
|
Raw
History Blame Contribute Delete
5.34 kB

Progress

2026-06-08

Completed

  • Confirmed Agent Tina is a FastAPI app, not a Streamlit app.
  • Added local setup helpers:
    • setup.ps1
    • run.ps1
  • Verified local dependencies, ffmpeg, app import, and local Uvicorn startup.
  • Switched default Whisper model from tiny to base.
  • Added Hugging Face Spaces Docker deployment support:
    • Dockerfile
    • Space metadata in README.md
    • deployment script at scripts/deploy_hf_space.py
  • Deployed Space:
    • Repo: https://huggingface.co/spaces/KPrashanth/agent-tina
    • Live app: https://kprashanth-agent-tina.hf.space/
  • Confirmed Hugging Face Space status: RUNNING.
  • Updated UI for the agreed meeting workflow:
    • meeting code
    • participant/device name label
    • meeting type
    • recording mode
    • known terms for future correction
    • same-room recording warning
    • raw transcript, corrected transcript, and MoM output sections
    • Markdown download buttons
  • Updated browser page title from Voice MoM Agent to Agent Tina.
  • Implemented the meeting-session MVP:
    • host-created meeting id and private host key
    • participant join links and host dashboard links
    • participant audio segments stored under meetings
    • Personal Mic and Room Recorder modes
    • room label and covered participants
    • LLM transcript correction for meeting segments
    • final MoM generation across all segments
    • server-side raw transcript, corrected transcript, and MoM Markdown files
    • host-only Markdown downloads
    • optional Hugging Face Dataset sync
  • Verified the deployed workflow end to end:
    • created a meeting through the live Space API
    • uploaded sample.wav
    • transcribed with Whisper
    • corrected transcript through OpenRouter
    • generated final MoM
    • downloaded minutes.md
    • confirmed metadata.json, raw transcript, corrected transcript, and MoM exist in the Dataset repo

Product Direction Agreed

  • Save raw transcripts, corrected transcripts, and MoM as Markdown files.
  • Add LLM-based transcript correction before MoM generation.
  • Support multi-device collaboration with participant name labels.
  • Do not start with automatic speaker detection or voice diarization.
  • For in-person meetings, prefer one Room Recorder device instead of every participant recording in the same room.
  • Final build plan documented in docs/final-build-plan.md.
  • Host-created meeting sessions will be the next major product step.

Next Implementation Target

Polish and harden the deployed workflow:

  • improve long-running transcription progress/error states
  • add meeting close/archive controls
  • show Hugging Face Dataset sync status
  • add regression tests for storage and API routes

2026-06-09

Completed

  • Evaluated data/recordings/day-end-meet-08-06-2026.aac.
  • Confirmed approximately 10-minute stereo AAC room recordings are supported.
  • Added context-guided Whisper transcription using meeting context, known terms, room label, and covered participants.
  • Increased default Whisper beam size to 5 for improved accuracy.
  • Added original audio preservation under each meeting folder.
  • Added optional original-audio Dataset sync.
  • Added host-only original-audio downloads.
  • Hardened host segment rendering against participant-supplied HTML.
  • Added transcript correction content-loss detection and raw-text fallback.
  • Separated correction and MoM models:
    • correction: openai/gpt-4o-mini
    • MoM: openai/gpt-oss-20b
  • Verified correction on the AAC sample with approximately 97.7% word retention in about 25 seconds.
  • Removed transcripts, outputs, and host keys from the public meeting endpoint.
  • Enforced private visibility for the meeting Dataset repo.
  • Increased meeting-id entropy to make shared join links harder to guess.
  • Verified the final deployed AAC workflow:
    • approximately 10-minute AAC Room Recorder segment processed in about 105 seconds
    • corrected transcript retained approximately 98% of raw transcript words
    • public meeting endpoint exposed no transcript segments
    • original 9.5 MB AAC downloaded successfully through host-only route
    • original audio, metadata, raw transcript, corrected transcript, and MoM persisted to the private Dataset repo
  • final MoM generation completed in about 8 seconds
  • Made long, multi-speaker Room Recorder audio the primary workflow.
  • Added a configurable 250 MB upload limit.
  • Added audio duration, format, channel, and sample-rate metadata.
  • Serialized Whisper transcription to protect CPU reliability under concurrent submissions.
  • Fixed stale-snapshot merging so simultaneous participant submissions do not overwrite each other.
  • Persisted original audio and uploaded status before long transcription begins.
  • Made in-person meetings and Room Recorder the UI defaults.
  • Required room label and covered participants for Room Recorder submissions.
  • Added atomic locked meeting storage for concurrent segment/status updates.
  • Added Dataset restore-on-demand after Space rebuilds.
  • Added lazy recovery so meeting metadata restores without downloading large audio until requested.
  • Added strict meeting-id path validation.
  • Verified deployed post-rebuild recovery:
    • older meeting metadata restored from the private Dataset
    • 9.5 MB original AAC restored lazily on host download
    • invalid meeting paths returned 404