grace-reader-study / README.md
DrSyedFaizan's picture
pin python 3.11 (fix audioop on 3.13)
464011c verified
|
Raw
History Blame Contribute Delete
3.65 kB

A newer version of the Gradio SDK is available: 6.23.1

Upgrade
metadata
title: GRACE Reader Study
emoji: 🩺
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 4.44.1
python_version: '3.11'
app_file: app.py
pinned: false

GRACE Reader Study annotation app

Blinded, resume-safe radiologist reader study for the GRACE project. One case per screen, no scrolling: reference CXR on the left, each anonymized system's output on the right with its rating controls directly beneath the image it refers to. Scale legends and subjective-term definitions are printed inline; per-image zoom / brightness / contrast are display-only.

What a reader does per case

For each anonymized item: answer correctness (Correct / Incorrect / Indeterminate), decision appropriateness (1-5), grounding relevance (Relevant / Partial / Not relevant), optional note. Plus one case-level question: is the reference (ground-truth) annotation acceptable. No best/worst ranking is asked; rankings are derived in the backend from per-item scores.

Run locally

pip install -r requirements.txt
python build_cases_example.py      # generates data/cases.json + placeholder images (demo)
python app.py                      # http://127.0.0.1:7860

Demo credentials (when READER_CREDENTIALS is unset): reader1 / changeme, reader2 / changeme.

Deploy as a Hugging Face Space

  1. Create a Gradio Space.
  2. Upload app.py, requirements.txt, README.md, and (for a demo) the data/ folder. For the real study do NOT commit patient images: put cases.json + images in a private dataset and set CASES_DATASET instead (see below).
  3. In Settings -> Variables and secrets, add these secrets (never commit them):
    • HF_TOKEN = your HF token with write permission (the value in Keys.txt; do not paste it into any file).
    • READER_CREDENTIALS = JSON, e.g. {"Dr Kotak":"<pw1>","Dr B":"<pw2>","Resident C":"<pw3>"}.
    • RESPONSE_DATASET = e.g. DrSyedFaizan/grace-reader-responses (created automatically, private).
    • CASES_DATASET (optional) = e.g. DrSyedFaizan/grace-reader-cases (private; pulled at boot).
    • APP_SECRET (optional) = any random string that signs resume tokens.

Data flow and storage

  • Responses stream to the private dataset RESPONSE_DATASET on every Save & Next, as responses/<reader>.jsonl, and to a local backup under local_responses/ (so an HF write hiccup never loses data).
  • Robust schema: one append-only record per (annotator, case_id, item_id) with a dims dict of value-per-dimension, plus a __case__ record for case-level answers. Later changes to layout or wording can never overwrite or invalidate saved annotations; reads take the latest record per key.

Auth / session / resume

  • First visit always lands on the login page; the app is gated by per-reader credentials.
  • On login the reader resumes at their first unfinished case and sees a live X / N counter (their own completed count, not a reset to 1). Item order is shuffled deterministically per (reader, case) so a resumed case looks identical.
  • Closing the tab does NOT log out (a signed token is kept in localStorage); use the explicit Logout button to end the session.

Preparing the real cases.json

Run build_cases_example.py and read its docstring for the exact schema. Each item's item_id is the TRUE system id (used only in the backend for the blinded GRACE-vs-baseline comparison and derived rankings); the reader never sees it. Keep items per case small (2-3) so the no-scroll goal holds. Enrich the case set per ../reader_protocol.md (near-boundary, deferrals, errors, rare pathology, easy anchors).