Spaces:
Sleeping
Sleeping
| title: Sokrates | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 4.44.0 | |
| python_version: "3.11" | |
| app_file: app.py | |
| pinned: false | |
| license: mit | |
| short_description: AI-assisted clinical intake assistant | |
| # Sokrates | |
| 🔗 **Social Media Post + Video:** https://x.com/guidoputignano/status/2065494688501506353 | |
| AI-assisted **clinical intake** for a medical visit. Sokrates listens to a | |
| doctor–patient conversation, transcribes it, automatically fills in a structured | |
| clinical intake form, and suggests the questions the doctor still needs to ask. | |
| Two sources of suggested questions: | |
| 1. **Gaps** — mandatory form fields that are still empty (deterministic engine). | |
| 2. **Clinical follow-ups** — coherent next questions generated by the LLM. | |
| > **Sokrates does not make diagnoses.** It prepares and structures data collection. | |
| ## Architecture | |
| ``` | |
| audio -> ASR (+ optional diarization) -> orchestration loop -> | |
| [structured extraction + gap engine + question generation] -> three-panel UI | |
| ``` | |
| - **UI**: Gradio Blocks, three columns — live transcript, the intake form filling | |
| in (green = filled, grey = missing), and a suggested-questions panel. | |
| - **ASR**: `faster-whisper` (large-v3), `language="en"`. Uploaded audio file | |
| (priority) and live microphone. Optional diarization via `pyannote` behind a flag. | |
| - **LLM**: any OpenAI-compatible endpoint (e.g. a vLLM server on Modal running | |
| **Qwen3-14B**, ≤ 32B params). Configured purely through environment variables. | |
| - **Schema**: oncology intake form as a Pydantic model (`sokrates/schema.py`). | |
| ## How to use | |
| 1. Provide a transcript one of three ways: | |
| - **Sample dialogue** — pick one and click *Load sample* (no audio needed). | |
| - **Upload / record a file** — click *Transcribe audio* to run ASR. | |
| - **Live microphone** — switch to the *Live microphone* tab and speak; | |
| Sokrates transcribes and updates the form automatically on pauses. | |
| 2. Click **Analyze transcript → update** (for the upload/sample paths). | |
| 3. Watch the **clinical form** fill in (green = filled, grey = optional missing, | |
| red-tinted = mandatory still missing) and the **suggested questions** appear | |
| (missing mandatory fields first, then coherent follow-ups). *Reset* clears the | |
| session. | |
| ## Modules | |
| | File | Role | | |
| | --------------------------- | ----------------------------------------------------------- | | |
| | `sokrates/schema.py` | Oncology intake form (Pydantic) + field metadata + merge. | | |
| | `sokrates/asr.py` | faster-whisper transcription (file + live mic), diarization.| | |
| | `sokrates/llm.py` | OpenAI-compatible client: extraction + question generation. | | |
| | `sokrates/gaps.py` | Deterministic gap engine (no LLM). | | |
| | `sokrates/orchestrator.py` | The loop: extract → gaps → questions. | | |
| | `app.py` | Gradio three-panel UI. | | |
| ## Local run | |
| ```bash | |
| python -m venv .venv && source .venv/bin/activate | |
| pip install -r requirements.txt | |
| python app.py | |
| ``` | |
| Open the local URL Gradio prints (default http://127.0.0.1:7860). | |
| ## Configuration (environment variables) | |
| No credentials live in the code. The LLM client is configured at runtime: | |
| | Variable | Purpose | Example | | |
| | ------------------- | ----------------------------------------------- | ---------------------------------------- | | |
| | `MODEL_BASE_URL` | OpenAI-compatible base URL | `https://your-modal-app.modal.run/v1` | | |
| | `MODEL_API_KEY` | API key / token for the endpoint | `sk-...` (or any token your server wants)| | |
| | `MODEL_NAME` | Model identifier served by the endpoint | `Qwen/Qwen3-14B` | | |
| | `ENABLE_DIARIZATION`| `1` to enable pyannote diarization (optional) | `0` | | |
| | `WHISPER_MODEL` | faster-whisper model size | `large-v3` | | |
| | `HF_TOKEN` | Hugging Face token (only if using pyannote) | `hf_...` | | |
| | `SOKRATES_GUIDED_JSON` | `1` to use vLLM `guided_json` instead of `response_format` | `0` | | |
| | `SOKRATES_NO_THINK` | `1` to disable Qwen3 thinking (faster, terser) | `0` | | |
| | `MODEL_TEMPERATURE` | sampling temperature (extraction uses `0.0`) | `0.0` | | |
| ### Pointing at a Modal vLLM endpoint | |
| This repo ships a ready-to-deploy Modal script (`modal_vllm.py`) that serves | |
| **Qwen3-14B** as an OpenAI-compatible vLLM endpoint. | |
| ```bash | |
| # 1. One-time setup | |
| pip install modal | |
| modal token new | |
| # Choose the API key the endpoint will require (reuse it as MODEL_API_KEY): | |
| modal secret create sokrates-llm MODEL_API_KEY=sk-sokrates-demo-123 | |
| # 2. Deploy (prints a https://...-serve.modal.run URL) | |
| modal deploy modal_vllm.py | |
| # 3. Point the app at it (note the trailing /v1) | |
| export MODEL_BASE_URL="https://<your-workspace>--sokrates-qwen3-14b-serve.modal.run/v1" | |
| export MODEL_API_KEY="sk-sokrates-demo-123" # must match the secret | |
| export MODEL_NAME="Qwen/Qwen3-14B" | |
| export SOKRATES_GUIDED_JSON=1 # vLLM enforces the JSON schema | |
| python app.py | |
| ``` | |
| Smoke-test the endpoint before launching the app: | |
| ```bash | |
| curl $MODEL_BASE_URL/models -H "Authorization: Bearer $MODEL_API_KEY" | |
| ``` | |
| No code changes are required to switch endpoints — only these variables. | |
| ## Deploy to a Hugging Face Space | |
| 1. Create a new **Gradio** Space. | |
| 2. Push this repo to the Space (or connect the GitHub repo). | |
| 3. In **Settings → Variables and secrets**, add `MODEL_BASE_URL`, `MODEL_API_KEY`, | |
| `MODEL_NAME` (and `HF_TOKEN` if using diarization) as **secrets**. | |
| 4. The Space runs `app.py` automatically. | |
| ## Sample dialogues | |
| Synthetic, English-only doctor–patient conversations live in `data/`: | |
| - `data/sample_dialogue_1.txt` — lung mass intake. | |
| - `data/sample_dialogue_2.txt` — breast lump intake. | |
| No real patient data is used anywhere in this project. | |