Spaces:
Sleeping
Sleeping
File size: 32,542 Bytes
d77360c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 | # Second Life β Project Reference (LLM Context Document)
DSCI 5260 | Group 7 | Last updated: 2026-04-25 (post Session 6 β inbox, messaging, trial dashboard, document upload, dataset seeding)
## Architecture Overview
Flask web app (port 5000) with two authenticated portals:
- **Patient Portal** (/patient): Register/login, update medical profile, get AI trial matches, connect with hospitals
- **Hospital Portal** (/hospital): Login, browse opt-in patients, search by condition, manage connections, edit profile
### Key Files
- `pipeline.py` β Core ML pipeline (data loading, feature engineering, model training, matching)
- `app.py` β Flask server: session auth, patient API, hospital API, pipeline API
- `database.py` β SQLite layer: 5 tables, auth, connections, trial interests, messages
- `templates/landing.html` β Login/register landing page
- `templates/patient.html` β Patient SPA (profile, trials, connections, inbox)
- `templates/hospital.html` β Hospital SPA (patients, search, my trials, inbox, connections, profile)
- `llm.md` β This file
## Running the System
```powershell
cd "E:\DSCI 5260\Project\PT"
python app.py
# open http://localhost:5000
```
On first run (no model_cache.pkl): loads all data files (~2-3 min), trains RF model, saves cache.
On subsequent runs: loads cached model immediately.
Delete `model_cache.pkl` to force retrain (required after pipeline feature changes).
## Demo Credentials
All hand-made patients have `open_to_trials=1` and password `pass123`.
| Username | Password | Name | Conditions | Location |
|----------|----------|------|-----------|----------|
| john_doe | pass123 | John Doe | hypertension, diabetes, MI | Boston, MA |
| jane_smith | pass123 | Jane Smith | asthma, atopic dermatitis, allergic rhinitis | Cambridge, MA |
| bob_jones | pass123 | Robert Jones | CAD, hypertension, chronic pain | Cleveland, OH |
| alice_brown | pass123 | Alice Brown | non-small cell lung cancer, stroke | Baltimore, MD |
| david_chen | pass123 | David Chen | diabetes, osteoporosis, CAD | Chicago, IL |
| Username | Password | Hospital | Location |
|----------|----------|---------|----------|
| mgh | mgh123 | Massachusetts General Hospital | Boston, MA |
| cleveland | clinic123 | Cleveland Clinic | Cleveland, OH |
| jhopkins | johns123 | Johns Hopkins Hospital | Baltimore, MD |
### Dataset-Backed Patients (Synthea)
20 real Synthea patients are auto-seeded on first run from `Final Patients Synthea Data/`. Username format: `synthea_<first 8 chars of Patient_ID>`, password `pass123`, all `open_to_trials=1`. Total DB patients: 25 (5 hand-made + 20 Synthea).
List all accounts: `SELECT username, first_name, last_name FROM patient_accounts ORDER BY username;`
---
## Data Sources
### Patient Side (Synthea synthetic data)
- `Final Patients Synthea Data/final_patients_conditions.csv` β ~967K rows, **265,893 patients**, 106 conditions overlapping with trials. Columns: Patient_ID, Condition_Name, Condition_End_Date
- `Final Patients Synthea Data/patients_details.csv` β Demographics. Columns: Patient_ID, First_Name, Last_Name, Birth_Date (DD-MM-YYYY), Gender, Race, Ethnicity, Address
- `Final Patients Synthea Data/patients_medications.csv` β **213,182 patients with med data**. Columns: Patient_ID, Medication_Name, Medication_End_Date
- `Final Patients Synthea Data/patients_observations.csv` β **23,231 patients with lab data**. Columns: Patient_ID, Observation_Name
### Trial Side (ClinicalTrials.gov / AACT)
- `Final Clinical Trails Data/trail_conditions.csv` β ~1M rows, **571,379 total trials**, **34,074 with matched condition profiles**. Columns: Trial_ID, Condition_Name_Lower
- `Final Clinical Trails Data/trail_eligibilities.csv` β Columns: Trial_ID, Gender (leading space β stripped), Minimum_Age, Maximum_Age, Eligibility_Criteria
- `Final Clinical Trails Data/trail_studies.csv` β **65,292 recruiting trials**. Columns: Trial_ID, Brief_Title, Overall_Status, Phase, Start_Date, Enrollment
- `Final Clinical Trails Data/trail_facilities.csv` β **189,274 trials with US state geo data**. Columns: Trial_ID, Facility_Name, Facility_City, Facility_State (full names), Facility_Country
- `Final Clinical Trails Data/trail_brief_summaries.csv` β Columns: Trial_ID, Brief_Summary
- `Final Clinical Trails Data/trail_interventions.csv` β **196,865 trials with drug intervention data**. Columns: Trial_ID, Intervention_Type, Intervention_Name
- `Final Clinical Trails Data/trail_countries.csv` β Not used for geo scoring (superseded by facility-level state data)
- `Final Clinical Trails Data/trail_keywords.csv` β Columns: Trial_ID, Keyword_Name_Lower
### MIMIC-IV Demo (code-level validation only, not in UI)
- `mimic-iv-clinical-database-demo-2.2/hosp/patients.csv.gz`
- `mimic-iv-clinical-database-demo-2.2/hosp/diagnoses_icd.csv.gz`
- `mimic-iv-clinical-database-demo-2.2/hosp/d_icd_diagnoses.csv.gz`
---
## SQLite Database (secondlife.db)
### Tables
```sql
patient_accounts(id, username, password_hash, synthea_id, first_name, last_name,
dob, gender, address, conditions TEXT DEFAULT '[]',
medications TEXT DEFAULT '[]', documents TEXT DEFAULT '[]',
open_to_trials INTEGER DEFAULT 0, created_at)
hospital_accounts(id, username, password_hash, hospital_name, location,
research_conditions TEXT DEFAULT '[]', created_at)
patient_trial_interests(id, patient_id, trial_id, trial_title, match_score,
status DEFAULT 'interested', created_at,
UNIQUE(patient_id, trial_id))
connections(id, patient_id, hospital_id, trial_id, trial_title,
initiated_by DEFAULT 'patient', status DEFAULT 'pending',
message, created_at,
UNIQUE(patient_id, hospital_id, trial_id))
-- Post-schema index handles NULL trial_id:
-- CREATE UNIQUE INDEX idx_conn_unique ON connections(patient_id, hospital_id, COALESCE(trial_id, ''))
connection_messages(id, connection_id, sender_role TEXT, -- 'patient' or 'hospital'
sender_id TEXT, body TEXT,
created_at, is_read INTEGER DEFAULT 0)
-- Index: idx_msgs_conn ON connection_messages(connection_id, created_at)
```
JSON fields (conditions, medications, documents, research_conditions) are stored as TEXT and parsed via `_row_to_dict()`.
### Key database.py Functions
| Function | Purpose |
|---------|---------|
| `get_open_patients_for_hospital(hid, condition_filter, include_connected)` | Returns open_to_trials=1 patients; when `include_connected=False` excludes already-connected patients |
| `get_connection_messages(connection_id)` | All messages for a connection, ASC order |
| `create_connection_message(connection_id, sender_role, sender_id, body)` | Insert message, returns dict |
| `mark_messages_read(connection_id, reader_role)` | Mark all messages from the other role as read |
| `unread_count(connection_id, reader_role)` | Count of unread messages from the other role |
| `get_hospital_inbox_threads(hospital_id)` | All threads with last_message, last_message_at, unread_count; sorted by activity |
| `get_patient_inbox_threads(patient_id)` | Same for patient side |
| `_seed_dataset_patients(c, max_patients=20)` | Seeds Synthea patients from CSV on first run |
---
## Feature Engineering (17 features in FEATURE_COLS)
| Feature | Description |
|---------|-------------|
| condition_overlap | Raw count of shared conditions |
| jaccard_similarity | overlap / union |
| overlap_ratio_trial | overlap / len(trial_conditions) |
| overlap_ratio_patient | overlap / len(patient_conditions) |
| condition_rarity_score | mean(1/log2(n_trials_per_cond+2)), normalised 0-1 |
| trial_specificity | 1 / trial_condition_count |
| condition_burden | total_patient_conds / 10 |
| active_ratio | active_conds / total_conds |
| resolved_ratio | resolved_conds / total_conds |
| age_distance | normalised distance outside age range (0 if within) |
| age_centered | position within age range (-1 to +1) |
| age_compatibility | 1.0 in range, decays over 30-year gap |
| gender_compatibility | 1.0 match/all, 0.1 mismatch |
| **geo_feasibility** | **State-level: 1.0 same state, 0.75 other US state, 0.5 no data** |
| **med_compatibility** | **Keyword overlap: patient meds vs trial drug interventions** |
| **lab_availability** | **Patient observation/lab type coverage (0-1, normalised by 20)** |
| data_completeness | fraction of key fields present |
---
## Flask API Routes
### Auth (no login required)
- `POST /auth/patient/register` β {success, patient} or {error}
- `POST /auth/patient/login` β {success, patient} or {error}
- `POST /auth/hospital/register` β {success, hospital} or {error}
- `POST /auth/hospital/login` β {success, hospital} or {error}
- `POST /auth/logout` β {success}
### Patient API (requires patient session)
- `GET /api/patient/profile` β patient dict (no password_hash)
- `POST /api/patient/profile` β updated patient dict; allowed fields: first_name, last_name, dob, gender, address, conditions, medications, open_to_trials
- `GET /api/patient/matches` β {results: [...trials], total}
- `GET /api/patient/interests` β {interests: [...]}
- `POST /api/patient/interest` β {success}; body: {trial_id, trial_title, match_score}
- `DELETE /api/patient/interest/<trial_id>` β {success}
- `GET /api/patient/connections` β {connections: [...]} joined with hospital_name
- `POST /api/patient/connect` β {success, connection} or 409 if duplicate; body: {hospital_id, trial_id, trial_title, message}
- `GET /api/patient/hospitals-for-trial?trial_id=NCT...` β {hospitals: [...]} tiered matching (see below)
- `GET /api/patient/connections/<cid>/messages` β {messages: [...]}; marks hospital messages read
- `POST /api/patient/connections/<cid>/messages` β {success, message}; body: {body}
- `GET /api/patient/inbox` β {threads: [...]} each with last_message, last_message_at, unread_count, hospital_name
- `GET /api/patient/documents` β {documents: [...]}
- `POST /api/patient/documents` β {success, document}; multipart/form-data file upload (max 10 MB, .pdf/.docx/.doc/.txt/.png/.jpg/.jpeg)
- `DELETE /api/patient/documents/<doc_id>` β {success}; removes file from disk and DB
- `GET /api/patient/documents/<doc_id>/download` β file download (as_attachment)
### Hospital API (requires hospital session)
- `GET /api/hospital/profile` β hospital dict (no password_hash)
- `POST /api/hospital/profile` β updated hospital dict; allowed fields: hospital_name, location, research_conditions
- `GET /api/hospital/patients?condition=&include_connected=` β {patients: [...]} open_to_trials=1; `include_connected=true` to include already-connected patients (used by Search tab)
- `POST /api/hospital/connect` β {success, connection} or 409 if duplicate; body: {patient_id, trial_id, trial_title, message}
- `GET /api/hospital/connections` β {connections: [...]} joined with patient fields
- `PUT /api/hospital/connections/<cid>/status` β {success}; body: {status: pending|accepted|rejected|completed}
- `GET /api/hospital/connections/<cid>/messages` β {messages: [...]}; marks patient messages read
- `POST /api/hospital/connections/<cid>/messages` β {success, message}; body: {body}
- `GET /api/hospital/inbox` β {threads: [...]} each with last_message, last_message_at, unread_count, first_name, last_name
- `GET /api/hospital/trials` β {trials: [...]} active trials matched to hospital profile (see below)
### Shared
- `GET /api/status` β {ready, stats} or {ready: false, message}
- `GET /api/conditions/autocomplete?q=...` β {results: [...]}
---
## Hospital Trial Dashboard (`/api/hospital/trials`)
`pipeline.trials_for_hospital(hospital_name, location, research_conditions, top_k=20)` β reverse of patient matching: given a hospital's profile, find active clinical trials it is most relevant to.
| Tier | Match condition | `match_reason` field |
|------|----------------|----------------------|
| 1 | Jaccard(hospital name tokens, trial facility name tokens) β₯ 0.25 | "name matched to trial site" |
| 2 | Hospital state matches a US trial facility state | "in same state as trial site" |
| 3 | Hospital `research_conditions` overlaps trial conditions | "researches related conditions" |
Active-only filter (`is_active` check) applied at every tier. Returns list of dicts:
`trial_id, title, phase, status, summary, location, facility_name, n_sites, match_tier, match_reason`
---
## Tiered Hospital Matching (`/api/patient/hospitals-for-trial`)
For each trial, hospitals in the DB are scored and returned in tier order (Tier 1 first):
| Tier | Match condition | `match_reason` field | UI label |
|------|----------------|----------------------|----------|
| 1 | Jaccard(hospital name tokens, any trial facility name tokens) β₯ 0.25 | "verified site on this trial" | Green β Verified Trial Sites |
| 2 | Hospital state (from "City, ST" location) matches a trial US facility state | "in same state as a trial site" | Grey β Related Hospitals |
| 3 | Hospital `research_conditions` overlaps trial conditions | "researches related conditions" | Grey β Related Hospitals |
| 4 | Fallback β trial has no facility/state data at all | "" | Grey β Related Hospitals |
The patient connect modal groups Tier 1 hospitals under a green "VERIFIED TRIAL SITES" header and Tiers 2-4 under a grey "RELATED HOSPITALS β not confirmed trial sites" header.
**Pipeline lookups used:**
- `pipeline.trial_facility_tokens[trial_id]` β list of frozensets of significant words from facility names
- `pipeline.trial_us_states[trial_id]` β set of US state full names (e.g. {"Massachusetts"})
- `pipeline.trial_profiles[trial_id]["conditions"]` β set of condition strings
**Stopword set for facility name tokenisation** (same in pipeline.py and app.py):
hospital, medical, center, centre, clinic, university, health, care, healthcare, system, institute, foundation, research, general, regional, national, community, services, department, division, college, school, the, of, and, at, for, in, a, an, is, by
---
## Trial Match Result Fields
Each item in `/api/patient/matches` results:
- trial_id, title, phase, status, min_age, max_age, sex, enrollment, start_date
- eligibility_probability (0-100, calibrated RF probability Γ 100)
- match_score (0-100, rule-based: overlap_ratio weighted)
- combined_score (0-100, 0.6 Γ eligibility + 0.4 Γ match_score)
- age_compatibility, gender_compatibility, geo_feasibility, med_compatibility (all 0-100)
- condition_rarity_score (0-1)
- overlap_conditions (list of conditions shared with patient)
- trial_conditions (all trial conditions)
- criteria (eligibility criteria text, truncated 500 chars)
- summary (brief summary, truncated 400 chars)
- **facility_name** (lead US facility name, or "" if not available)
- location (lead US facility city/state/country string)
- n_sites (total facility count for this trial)
- interest_status (null | 'interested' | 'withdrawn', from patient_trial_interests)
---
## Data Privacy Model
1. Hospitals see only patients with open_to_trials=1 (name, age, gender, conditions)
2. Full details accessible only after patient-initiated connection
3. Hospital cannot contact a patient unless patient is open to trials
4. Connection record: patient_id, hospital_id, trial_id, initiated_by, status, message
5. Hospital can also initiate connections with opt-in patients from the hospital portal
---
## XSS Prevention
All user-controlled strings use DOM API (never innerHTML for user data):
```javascript
function escH(s) { // text content in innerHTML contexts
const d = document.createElement('div');
d.appendChild(document.createTextNode(String(s||'')));
return d.innerHTML;
}
function escA(s) { // HTML attribute values
return String(s||'').replace(/&/g,'&').replace(/"/g,'"')
.replace(/</g,'<').replace(/>/g,'>');
}
```
Event listeners use addEventListener only. Tags and cards built via createElement + textContent.
---
## ML Model
- Random Forest (n_estimators=200, max_depth=12, class_weight="balanced")
- CalibratedClassifierCV (isotonic, cv=3) for probability calibration
- GroupShuffleSplit (patient-level, 80/20, no leakage) for train/test split
- GroupKFold (5-fold, patient-level) for cross-validation
- Training sample: 3000 patients Γ 30 trials each + random negatives
- Cache: `model_cache.pkl` (delete to force retrain)
### Actual Metrics (from verified live run, 2026-04-25)
| Metric | Value |
|--------|-------|
| Accuracy | 85.15% |
| AUC-ROC | 0.5976 |
| CV AUC (5-fold) | 0.5992 Β± 0.0052 |
| F1 | 0.9168 |
| Precision | 0.8518 |
| Recall | 0.9925 |
| Brier score | 0.1263 |
| Avg precision | 0.8540 |
| Train size | 90,994 pairs |
| Test size | 22,681 pairs |
| Positive label rate | 82.1% |
> **Note on AUC:** The 82.1% positive rate in pseudo-labels (weighted 6-feature labelling threshold at 0.5) makes the classification task easy to solve trivially β high accuracy/recall but lower AUC. To improve AUC, the pseudo-label threshold should be raised (e.g. 0.6) or positive/negative sampling balanced more aggressively.
### Feature Importance (Random Forest, ranked)
| Rank | Feature | Importance |
|------|---------|-----------|
| 1 | age_distance | 0.2040 |
| 2 | age_compatibility | 0.1691 |
| 3 | gender_compatibility | 0.1288 |
| 4 | age_centered | 0.0791 |
| 5 | jaccard_similarity | 0.0751 |
| 6 | condition_rarity_score | 0.0621 |
| 7 | overlap_ratio_trial | 0.0519 |
| 8 | overlap_ratio_patient | 0.0512 |
| 9 | condition_overlap | 0.0396 |
| 10 | condition_burden | 0.0274 |
| 11 | resolved_ratio | 0.0263 |
| 12 | active_ratio | 0.0258 |
| 13 | lab_availability | 0.0208 |
| 14 | geo_feasibility | 0.0158 |
| 15 | med_compatibility | 0.0144 |
| 16 | trial_specificity | 0.0087 |
| 17 | data_completeness | 0.0000 |
---
## MIMIC-IV Validation (code-level only, not in UI)
- 100 demo patients, ~90% match rate after 3-tier ICD β condition mapping
- Call: `pipeline.validate_mimic()` β list of {subject_id, mapped_conditions, n_matches, top_match}
- 3-tier mapping: exact β substring containment β word-overlap β₯ 75%
- Not exposed via any Flask route
---
## All Bug Fixes by Session
### Session 3 Fixes (2026-04-25) β Two-portal foundation
#### Fix 1 β patient_id not passed to match_patient() (CRITICAL)
**Before:** `pipeline.match_patient(conditions, age, gender, top_k=20)`
**After:** `pipeline.match_patient(conditions, age, gender, top_k=20, patient_id=session["patient_id"], address=address)`
Without this, patient-specific medication keywords and lab scores defaulted to empty / 0.3 for all users β med_compatibility and lab_availability were effectively constants.
#### Fix 2 β Pseudo-label used only 3 features (HIGH)
**Before:** AND gate on age/gender/condition β geo/med/lab had near-zero training influence.
**After:** 6-feature weighted score with 15% random noise:
```python
score = (
0.30 * float(row["age_compatibility"] > 0.6) +
0.15 * float(row["gender_compatibility"] > 0.5) +
0.25 * float(row["jaccard_similarity"] > 0.05) +
0.10 * float(row["geo_feasibility"]) +
0.10 * float(row["med_compatibility"]) +
0.10 * float(row["lab_availability"])
)
base = int(score >= 0.5)
# 15% hash-deterministic noise for realism
```
#### Fix 3 β geo_feasibility was country-level heuristic (MEDIUM)
**Before:** Float from `trail_countries.csv` (1.0 US, 0.7 multi-national, 0.35 non-US). No patient location.
**After:** State-level matching using `trail_facilities.csv` + patient address regex:
```python
def _geo_score(patient_state_full, trial_states):
if not trial_states: return 0.5 # no US facility data β neutral
if patient_state_full in trial_states: return 1.0
return 0.75 # other US state
```
#### Fix 4 β NameError `trial_geo` in _compute_features return dict
`"geo_feasibility": float(trial_geo)` β `"geo_feasibility": geo_feasibility`
#### Fix 5 β XSS in condition tag onclick handlers (MEDIUM)
`addConditionTag('${c}')` broke for conditions with apostrophes (e.g. "alzheimer's disease").
Fixed with DOM-based `makeTag()` using textContent + addEventListener. No inline onclick anywhere.
#### Fix 6 β Duplicate connection prevention (was: no guard)
- `connections` table: added `UNIQUE(patient_id, hospital_id, trial_id)` schema constraint
- `init_db()`: runs `CREATE UNIQUE INDEX IF NOT EXISTS idx_conn_unique ON connections(patient_id, hospital_id, COALESCE(trial_id, ''))` to handle NULL trial_id and backfill existing DBs
- `create_connection()`: pre-checks `connection_exists()` before insert; returns `None` on duplicate
- `/api/patient/connect` and `/api/hospital/connect`: return 409 when `create_connection()` returns None
#### Fix 7 β Hospital patient feed showed already-contacted patients (was: no exclusion)
`get_open_patients_for_hospital()` now uses:
```sql
WHERE open_to_trials=1
AND id NOT IN (SELECT DISTINCT patient_id FROM connections WHERE hospital_id=?)
```
#### Fix 8 β Demo seed: john_doe starts with open_to_trials=1
Hospital portal was empty on a fresh database. `_seed_demo_data()` now seeds john_doe with `open_to_trials=1`.
#### Fix 9 β Hospital registration silently ignored research_conditions
`templates/landing.html` hospital register form now collects comma-separated research conditions and sends them as a parsed lowercase array to the backend.
#### Fix 10 β Trial cards only showed site count, not facility name or location
`pipeline.py match_patient()` now extracts `facility_name` from `Facility_Name` column; prefers US facilities. Patient portal detail grid shows "Lead Site" and "Location" when available.
---
### Session 4 Fixes (2026-04-25) β Hospital matching overhaul + profile editing
#### Fix 11 β Hospital suggestion logic replaced (was: research_conditions overlap only)
Complete replacement of `/api/patient/hospitals-for-trial`:
**Before:** looped all hospitals, included any whose `research_conditions` overlapped trial conditions. No tier concept, no facility data used.
**After:** 4-tier system using two new pipeline lookups:
- `pipeline.trial_facility_tokens[trial_id]` β built from `Facility_Name` column in trail_facilities.csv, US rows only. Each facility name tokenised by stripping stopwords + words < 3 chars.
- `pipeline.trial_us_states[trial_id]` β set of full US state names for the trial
Helper functions in `app.py`:
```python
def _hospital_name_tokens(name: str) -> frozenset:
# strips stopwords, keeps words β₯ 3 chars
...
def _facility_match_score(h_tokens: frozenset, facility_token_list: list) -> float:
# best Jaccard score against any facility in the trial
...
```
Each hospital gets one tier assigned and a `match_reason` + `match_tier` in the response.
Result list sorted by `match_tier` ascending (best first).
#### Fix 12 β Hospital portal had no profile editing
`POST /api/hospital/profile` added (was GET-only). `database.py update_hospital_profile()` added. `templates/hospital.html` now has a **My Profile** tab with editable hospital name, location, and research condition tags. On save, the navbar hospital name updates live without a page reload.
#### Fix 13 β Model disclaimer missing from patient trial results
`templates/patient.html` trial results section now shows an alert above results:
> "Match percentages are predictions from a model trained on synthetic patient data and rule-based labels β not validated clinical eligibility determinations. Always consult a healthcare provider before enrolling in any trial."
---
### Session 5 Fixes (2026-04-25) β Modal UX + bug fixes
#### Fix 14 β Patient connect modal showed all hospitals in one flat list
**Before:** All hospitals (all tiers) in a single flat list, sorted by tier, with coloured badges as the only visual distinction.
**After:** Modal renders two visually separated sections:
- **"VERIFIED TRIAL SITES"** (green `sec-head`) β Tier 1 hospitals only
- **"RELATED HOSPITALS β not confirmed trial sites"** (grey `sec-head` with inline subtitle) β Tiers 2, 3, 4
Both sections only render if they have entries. Click delegation on the outer `#hospitalList` wrapper still works for both sections.
#### Fix 15 β Close button invisible on hospital profile condition tags
`hospital.html renderProfTags()`: `btn-close-white` (white X) on `bg-info text-dark` badge (light blue background) β `btn-close` (dark X). The X was invisible before.
#### Fix 16 β Login/register forms required mouse click, no Enter key support
`templates/landing.html`: Added `_onEnter(inputId, fn)` helper and wired Enter key on all login and register inputs (both patient and hospital portals). Works on username field too (not just password).
#### Fix 17 β System ready banner never auto-cleared on slow boot
`templates/patient.html`: `checkStatus()` was called once at DOMContentLoaded and never again. If the pipeline was still training when the user opened the page, the yellow banner persisted even after the pipeline finished.
**After:** `startStatusPoll()` starts a `setInterval` (5s) when the initial check finds `ready: false`. The interval clears itself once `ready: true` is received.
```javascript
checkStatus().then(() => { if (!sysReady) startStatusPoll(); });
```
---
## Live Run Verification (2026-04-25)
End-to-end test results after full retrain with no model_cache.pkl:
| Test | Result |
|------|--------|
| Landing page GET / | 200 OK |
| Patient login john_doe/pass123 | OK β returns patient JSON |
| Hospital login mgh/mgh123 | OK β returns hospital JSON |
| Pipeline ready (api/status) | ready: true |
| /api/patient/matches for john_doe | 20 results, all score fields populated |
| Top match geo score for Greece trial | 50% (no US facility β correct) |
| Hospital browses open_to_trials patients | 25 patients visible (5 demo + 20 Synthea) |
| Hospital condition search ?condition=hypertension | results including Synthea patients |
| Hospital β patient connect (POST) | OK, status=pending |
| Patient sees hospital connection (GET) | 1 connection, hospital_name present |
| /api/conditions/autocomplete?q=hyper | ["hypertension"] |
| Duplicate connect attempt | 409 error |
| Hospital profile save | navbar name updates live |
| Patient connect modal | Two sections render correctly |
| /api/hospital/trials for mgh | Active trials with match tiers |
| /api/hospital/inbox | Threads with unread counts |
| /api/patient/inbox | Threads with hospital names |
| Patient document upload | File saved, metadata in DB |
| Inbox Synthea seeding | 25 patients total confirmed |
---
### Session 6 Fixes (2026-04-25) β Messaging, inbox, trial dashboard, document upload, dataset seeding
#### Fix 18 β Hospital trial dashboard (My Trials tab)
**Before:** Hospital portal had no way to see which clinical trials were relevant to it.
**After:** New "My Trials" nav tab in `hospital.html`. Calls `GET /api/hospital/trials` β `pipeline.trials_for_hospital()`. Active-only filter at all 3 tiers. Cards show status badge, phase, match tier, facility name, summary excerpt, and a "View on ClinicalTrials.gov β" link.
Active-only filter: `if not self.trial_profiles[trial_id].get("is_active", False): continue` at each tier loop in `pipeline.py`.
#### Fix 19 β Messaging (chat in connections)
**Before:** Connections table had no messaging. Patients and hospitals could only see connection status.
**After:**
- New `connection_messages` table with `sender_role`, `sender_id`, `body`, `is_read`.
- `GET/POST /api/patient/connections/<cid>/messages` and `GET/POST /api/hospital/connections/<cid>/messages`.
- Both portals have a messages modal (`#msgModal`) opened by a Chat button in the Connections table.
- `mark_messages_read()` called on GET to auto-mark messages as read when the recipient opens the thread.
#### Fix 20 β Dedicated Inbox tab (both portals)
**Before:** Chat only accessible from the My Connections table row β no inbox overview.
**After:** New "Inbox" nav tab in both `hospital.html` and `patient.html`.
- Calls `GET /api/hospital/inbox` or `GET /api/patient/inbox`.
- Backed by `get_hospital_inbox_threads()` / `get_patient_inbox_threads()` β SQL subqueries aggregate last_message, last_message_at, unread_count per thread.
- Threads sorted by most recent activity (Python-side sort on `last_message_at or created_at`).
- Unread count badge on nav tab button updates when inbox loads.
- "Open" button reuses the existing `openMsgModal()` and messages modal.
#### Fix 21 β Document upload (patient portal)
**Before:** Patient profile had no file upload section.
**After:** "My Documents" card added to patient profile tab. 4 routes:
- `POST /api/patient/documents` β werkzeug `secure_filename`, 10 MB limit, allowed extensions: `.pdf/.docx/.doc/.txt/.png/.jpg/.jpeg`. Saves to `uploads/patient_docs/<patient_id>/`. Metadata stored as JSON array in `patient_accounts.documents`.
- `DELETE /api/patient/documents/<doc_id>` β removes file from disk and metadata from DB.
- `GET /api/patient/documents/<doc_id>/download` β serves file as attachment.
- `app.config["MAX_CONTENT_LENGTH"] = 10 * 1024 * 1024` enforced Flask-side.
#### Fix 22 β Dataset-backed patient seeding
**Before:** Only hand-made demo patients in DB (john_doe only had open_to_trials=1 initially). Hospital search returned 0 results on fresh DB.
**After:** `_seed_dataset_patients(c, max_patients=20)` in `database.py` seeds 20 real Synthea patients from CSV files on first run. Skips deceased patients (Death_Date not empty). Reads up to 8 conditions + 6 medications per patient. Username: `synthea_<first8chars_of_Patient_ID>`, password `pass123`, `open_to_trials=1`.
All 5 hand-made demo patients also set to `open_to_trials=1`. Total: 25 patients in DB.
#### Fix 23 β Hospital search include_connected toggle
**Before:** Hospital Search tab also excluded already-connected patients, same as Available Patients tab β making it useless for re-searching.
**After:** Search tab adds `include_connected=true` query param. `GET /api/hospital/patients?include_connected=true` bypasses the exclusion subquery. Available Patients tab retains strict exclusion. Toggle checkbox in Search tab UI.
#### Fix 24 β Template auto-reload
**Before:** `app.run(debug=False, use_reloader=False)` β template edits required server restart to take effect.
**After:** `app.config["TEMPLATES_AUTO_RELOAD"] = True` added after other config lines. Templates now reload on every request without enabling full debug mode or the reloader.
---
## Known Issues / Future Improvements
1. **High pseudo-label positive rate (82.1%)** β lowers AUC-ROC to ~0.60. Fix: raise label threshold from 0.5 to 0.6, or explicitly sample equal positive/negative pairs.
2. **data_completeness feature importance = 0** β nearly constant across training pairs (all synthetic patients have complete data). Consider removing from FEATURE_COLS.
3. **Fuzzy condition matching not implemented** β only exact condition name overlaps used (106 conditions). Substring/semantic fuzzy matching would expand coverage significantly.
4. **lab_availability coverage is low (8.7%)** β observations file is sparse. Consider normalising denominator to the subset that has any lab data.
5. **Hospital portal does not rank patients by match quality** β listed in insertion order. Could rank by condition overlap with the hospital's research_conditions.
6. **No email/notification system** β connection requests visible only inside the portal.
7. **No automated tests** β syntax checking only (`python -m py_compile`). Key flows to cover: patient registration β profile update β match β connect; hospital registration β patient browse β connect β status update; duplicate connection rejection.
8. **Inbox badge not auto-refreshed** β unread count badge only updates when the user clicks the Inbox tab. No real-time push; would require polling or WebSockets.
9. **Document access control** β uploaded files are served from disk by doc_id only; no additional hospital-side access to patient documents (by design β privacy model). Hospital sees document count in patient profile only after connection.
10. **Synthea patients have Synthea-style names** (e.g. "Geovany567 Reichert456") β cosmetically odd but functionally correct. No fix needed for demo.
|