| --- |
| title: GRACE Reader Study |
| emoji: 🩺 |
| colorFrom: indigo |
| colorTo: blue |
| sdk: gradio |
| sdk_version: 4.44.1 |
| python_version: "3.11" |
| app_file: app.py |
| pinned: false |
| --- |
| |
| # GRACE Reader Study annotation app |
|
|
| Blinded, resume-safe radiologist reader study for the GRACE project. One case per screen, |
| no scrolling: reference CXR on the left, each anonymized system's output on the right with its |
| rating controls directly beneath the image it refers to. Scale legends and subjective-term |
| definitions are printed inline; per-image zoom / brightness / contrast are display-only. |
|
|
| ## What a reader does per case |
| For each anonymized item: **answer correctness** (Correct / Incorrect / Indeterminate), |
| **decision appropriateness** (1-5), **grounding relevance** (Relevant / Partial / Not relevant), |
| optional note. Plus one case-level question: is the reference (ground-truth) annotation acceptable. |
| No best/worst ranking is asked; rankings are derived in the backend from per-item scores. |
|
|
| ## Run locally |
| ```bash |
| pip install -r requirements.txt |
| python build_cases_example.py # generates data/cases.json + placeholder images (demo) |
| python app.py # http://127.0.0.1:7860 |
| ``` |
| Demo credentials (when READER_CREDENTIALS is unset): `reader1 / changeme`, `reader2 / changeme`. |
| |
| ## Deploy as a Hugging Face Space |
| 1. Create a **Gradio** Space. |
| 2. Upload `app.py`, `requirements.txt`, `README.md`, and (for a demo) the `data/` folder. |
| For the real study do NOT commit patient images: put `cases.json` + images in a **private |
| dataset** and set `CASES_DATASET` instead (see below). |
| 3. In **Settings -> Variables and secrets**, add these **secrets** (never commit them): |
| - `HF_TOKEN` = your HF token with **write** permission (the value in `Keys.txt`; do not paste it into any file). |
| - `READER_CREDENTIALS` = JSON, e.g. `{"Dr Kotak":"<pw1>","Dr B":"<pw2>","Resident C":"<pw3>"}`. |
| - `RESPONSE_DATASET` = e.g. `DrSyedFaizan/grace-reader-responses` (created automatically, private). |
| - `CASES_DATASET` (optional) = e.g. `DrSyedFaizan/grace-reader-cases` (private; pulled at boot). |
| - `APP_SECRET` (optional) = any random string that signs resume tokens. |
|
|
| ## Data flow and storage |
| - Responses stream to the private dataset `RESPONSE_DATASET` on every **Save & Next**, as |
| `responses/<reader>.jsonl`, and to a local backup under `local_responses/` (so an HF write |
| hiccup never loses data). |
| - **Robust schema:** one append-only record per `(annotator, case_id, item_id)` with a |
| `dims` dict of value-per-dimension, plus a `__case__` record for case-level answers. Later |
| changes to layout or wording can never overwrite or invalidate saved annotations; reads take |
| the latest record per key. |
|
|
| ## Auth / session / resume |
| - First visit always lands on the login page; the app is gated by per-reader credentials. |
| - On login the reader resumes at their first unfinished case and sees a live `X / N` counter |
| (their own completed count, not a reset to 1). Item order is shuffled deterministically per |
| `(reader, case)` so a resumed case looks identical. |
| - Closing the tab does NOT log out (a signed token is kept in `localStorage`); use the explicit |
| **Logout** button to end the session. |
|
|
| ## Preparing the real cases.json |
| Run `build_cases_example.py` and read its docstring for the exact schema. Each item's `item_id` |
| is the TRUE system id (used only in the backend for the blinded GRACE-vs-baseline comparison and |
| derived rankings); the reader never sees it. Keep `items` per case small (2-3) so the no-scroll |
| goal holds. Enrich the case set per `../reader_protocol.md` (near-boundary, deferrals, errors, |
| rare pathology, easy anchors). |
|
|