backyard-radiology-professor / docs /deploy_modal_backend.md
imadreamerboy's picture
Configure Modal-backed Space proxy
ee7c461 verified
|
Raw
History Blame Contribute Delete
2.72 kB

Deploy the Modal GPU backend

The official Build Small Space must live under build-small-hackathon/backyard-radiology-professor. Hugging Face rejected paid L4 hardware for that org without billing credits, and ZeroGPU is not available for Docker Spaces. The production path is therefore:

  1. Run the real backend on Modal GPU.
  2. Configure the official Hugging Face Space as a Gradio proxy with RAD_TRAINER_REMOTE_BACKEND_URL.

Modal deployment:

uv sync --extra deploy
$env:HF_TOKEN = (wsl.exe -d Ubuntu -- bash -lc 'cat ~/.cache/huggingface/token').Trim()
uv run modal setup
uv run modal deploy deploy/modal_app.py

The Modal app uses:

  • Dockerfile.modal, equivalent to the local CUDA runtime without BuildKit-only cache mounts.
  • An L40S GPU profile for full MedGemma 27B + 4B residency.
  • A Modal Volume mounted at /data for GGUF, X-Raydar, Hugging Face caches, and transient session state.
  • HF_TOKEN from the local environment during deploy, or a Modal secret named backyard-radiology-professor-hf.
  • Modal proxy authentication. Requests without the private key and secret are rejected before a GPU container starts.
  • min_containers=0 and a 60-second scale-down window. The app remains deployed but has no GPU compute cost after its last request finishes and the warm window expires.

Create a proxy-auth token in the Modal dashboard as documented in Webhook proxy auth. Store its key and secret locally; do not commit them.

$env:RAD_TRAINER_MODAL_KEY = "wk-..."
$env:RAD_TRAINER_MODAL_SECRET = "ws-..."

Deploy Modal, copy its web URL, then configure the official Space. The script stores the credentials as Hugging Face Space secrets:

uv run modal deploy deploy/modal_app.py
uv run python scripts/configure_hf_space.py `
  --backend-url https://YOUR-MODAL-URL.modal.run

Then validate the official Space URL:

uv run python scripts/validate_golden_cases.py --app-url https://build-small-hackathon-backyard-radiology-professor.hf.space --timeout-seconds 1800
uv run python scripts/benchmark_runtime.py --app-url https://build-small-hackathon-backyard-radiology-professor.hf.space --runs 2 --timeout-seconds 1800 --gpu-baseline-mb 0

The Space serves its bundled case catalog and status locally, so opening the page does not wake Modal. Opening or analyzing a study starts the backend.

Check the current deployment without waking it:

uv run modal app list
uv run modal container list

Leave the Modal app deployed for ordinary idling. modal app stop permanently stops the deployment and requires another modal deploy; it is not the normal cost-control mechanism.