--- title: GRACE Reader Study emoji: 🩺 colorFrom: indigo colorTo: blue sdk: gradio sdk_version: 4.44.1 python_version: "3.11" app_file: app.py pinned: false --- # GRACE Reader Study annotation app Blinded, resume-safe radiologist reader study for the GRACE project. One case per screen, no scrolling: reference CXR on the left, each anonymized system's output on the right with its rating controls directly beneath the image it refers to. Scale legends and subjective-term definitions are printed inline; per-image zoom / brightness / contrast are display-only. ## What a reader does per case For each anonymized item: **answer correctness** (Correct / Incorrect / Indeterminate), **decision appropriateness** (1-5), **grounding relevance** (Relevant / Partial / Not relevant), optional note. Plus one case-level question: is the reference (ground-truth) annotation acceptable. No best/worst ranking is asked; rankings are derived in the backend from per-item scores. ## Run locally ```bash pip install -r requirements.txt python build_cases_example.py # generates data/cases.json + placeholder images (demo) python app.py # http://127.0.0.1:7860 ``` Demo credentials (when READER_CREDENTIALS is unset): `reader1 / changeme`, `reader2 / changeme`. ## Deploy as a Hugging Face Space 1. Create a **Gradio** Space. 2. Upload `app.py`, `requirements.txt`, `README.md`, and (for a demo) the `data/` folder. For the real study do NOT commit patient images: put `cases.json` + images in a **private dataset** and set `CASES_DATASET` instead (see below). 3. In **Settings -> Variables and secrets**, add these **secrets** (never commit them): - `HF_TOKEN` = your HF token with **write** permission (the value in `Keys.txt`; do not paste it into any file). - `READER_CREDENTIALS` = JSON, e.g. `{"Dr Kotak":"","Dr B":"","Resident C":""}`. - `RESPONSE_DATASET` = e.g. `DrSyedFaizan/grace-reader-responses` (created automatically, private). - `CASES_DATASET` (optional) = e.g. `DrSyedFaizan/grace-reader-cases` (private; pulled at boot). - `APP_SECRET` (optional) = any random string that signs resume tokens. ## Data flow and storage - Responses stream to the private dataset `RESPONSE_DATASET` on every **Save & Next**, as `responses/.jsonl`, and to a local backup under `local_responses/` (so an HF write hiccup never loses data). - **Robust schema:** one append-only record per `(annotator, case_id, item_id)` with a `dims` dict of value-per-dimension, plus a `__case__` record for case-level answers. Later changes to layout or wording can never overwrite or invalidate saved annotations; reads take the latest record per key. ## Auth / session / resume - First visit always lands on the login page; the app is gated by per-reader credentials. - On login the reader resumes at their first unfinished case and sees a live `X / N` counter (their own completed count, not a reset to 1). Item order is shuffled deterministically per `(reader, case)` so a resumed case looks identical. - Closing the tab does NOT log out (a signed token is kept in `localStorage`); use the explicit **Logout** button to end the session. ## Preparing the real cases.json Run `build_cases_example.py` and read its docstring for the exact schema. Each item's `item_id` is the TRUE system id (used only in the backend for the blinded GRACE-vs-baseline comparison and derived rankings); the reader never sees it. Keep `items` per case small (2-3) so the no-scroll goal holds. Enrich the case set per `../reader_protocol.md` (near-boundary, deferrals, errors, rare pathology, easy anchors).