File size: 32,542 Bytes
d77360c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
# Second Life β€” Project Reference (LLM Context Document)
DSCI 5260 | Group 7 | Last updated: 2026-04-25 (post Session 6 β€” inbox, messaging, trial dashboard, document upload, dataset seeding)

## Architecture Overview

Flask web app (port 5000) with two authenticated portals:
- **Patient Portal** (/patient): Register/login, update medical profile, get AI trial matches, connect with hospitals
- **Hospital Portal** (/hospital): Login, browse opt-in patients, search by condition, manage connections, edit profile

### Key Files
- `pipeline.py` β€” Core ML pipeline (data loading, feature engineering, model training, matching)
- `app.py` β€” Flask server: session auth, patient API, hospital API, pipeline API
- `database.py` β€” SQLite layer: 5 tables, auth, connections, trial interests, messages
- `templates/landing.html` β€” Login/register landing page
- `templates/patient.html` β€” Patient SPA (profile, trials, connections, inbox)
- `templates/hospital.html` β€” Hospital SPA (patients, search, my trials, inbox, connections, profile)
- `llm.md` β€” This file

## Running the System

```powershell
cd "E:\DSCI 5260\Project\PT"
python app.py
# open http://localhost:5000
```

On first run (no model_cache.pkl): loads all data files (~2-3 min), trains RF model, saves cache.
On subsequent runs: loads cached model immediately.

Delete `model_cache.pkl` to force retrain (required after pipeline feature changes).

## Demo Credentials

All hand-made patients have `open_to_trials=1` and password `pass123`.

| Username | Password | Name | Conditions | Location |
|----------|----------|------|-----------|----------|
| john_doe | pass123 | John Doe | hypertension, diabetes, MI | Boston, MA |
| jane_smith | pass123 | Jane Smith | asthma, atopic dermatitis, allergic rhinitis | Cambridge, MA |
| bob_jones | pass123 | Robert Jones | CAD, hypertension, chronic pain | Cleveland, OH |
| alice_brown | pass123 | Alice Brown | non-small cell lung cancer, stroke | Baltimore, MD |
| david_chen | pass123 | David Chen | diabetes, osteoporosis, CAD | Chicago, IL |

| Username | Password | Hospital | Location |
|----------|----------|---------|----------|
| mgh | mgh123 | Massachusetts General Hospital | Boston, MA |
| cleveland | clinic123 | Cleveland Clinic | Cleveland, OH |
| jhopkins | johns123 | Johns Hopkins Hospital | Baltimore, MD |

### Dataset-Backed Patients (Synthea)
20 real Synthea patients are auto-seeded on first run from `Final Patients Synthea Data/`. Username format: `synthea_<first 8 chars of Patient_ID>`, password `pass123`, all `open_to_trials=1`. Total DB patients: 25 (5 hand-made + 20 Synthea).

List all accounts: `SELECT username, first_name, last_name FROM patient_accounts ORDER BY username;`

---

## Data Sources

### Patient Side (Synthea synthetic data)
- `Final Patients Synthea Data/final_patients_conditions.csv` β€” ~967K rows, **265,893 patients**, 106 conditions overlapping with trials. Columns: Patient_ID, Condition_Name, Condition_End_Date
- `Final Patients Synthea Data/patients_details.csv` β€” Demographics. Columns: Patient_ID, First_Name, Last_Name, Birth_Date (DD-MM-YYYY), Gender, Race, Ethnicity, Address
- `Final Patients Synthea Data/patients_medications.csv` β€” **213,182 patients with med data**. Columns: Patient_ID, Medication_Name, Medication_End_Date
- `Final Patients Synthea Data/patients_observations.csv` β€” **23,231 patients with lab data**. Columns: Patient_ID, Observation_Name

### Trial Side (ClinicalTrials.gov / AACT)
- `Final Clinical Trails Data/trail_conditions.csv` β€” ~1M rows, **571,379 total trials**, **34,074 with matched condition profiles**. Columns: Trial_ID, Condition_Name_Lower
- `Final Clinical Trails Data/trail_eligibilities.csv` β€” Columns: Trial_ID, Gender (leading space β€” stripped), Minimum_Age, Maximum_Age, Eligibility_Criteria
- `Final Clinical Trails Data/trail_studies.csv` β€” **65,292 recruiting trials**. Columns: Trial_ID, Brief_Title, Overall_Status, Phase, Start_Date, Enrollment
- `Final Clinical Trails Data/trail_facilities.csv` β€” **189,274 trials with US state geo data**. Columns: Trial_ID, Facility_Name, Facility_City, Facility_State (full names), Facility_Country
- `Final Clinical Trails Data/trail_brief_summaries.csv` β€” Columns: Trial_ID, Brief_Summary
- `Final Clinical Trails Data/trail_interventions.csv` β€” **196,865 trials with drug intervention data**. Columns: Trial_ID, Intervention_Type, Intervention_Name
- `Final Clinical Trails Data/trail_countries.csv` β€” Not used for geo scoring (superseded by facility-level state data)
- `Final Clinical Trails Data/trail_keywords.csv` β€” Columns: Trial_ID, Keyword_Name_Lower

### MIMIC-IV Demo (code-level validation only, not in UI)
- `mimic-iv-clinical-database-demo-2.2/hosp/patients.csv.gz`
- `mimic-iv-clinical-database-demo-2.2/hosp/diagnoses_icd.csv.gz`
- `mimic-iv-clinical-database-demo-2.2/hosp/d_icd_diagnoses.csv.gz`

---

## SQLite Database (secondlife.db)

### Tables
```sql
patient_accounts(id, username, password_hash, synthea_id, first_name, last_name,
                 dob, gender, address, conditions TEXT DEFAULT '[]',
                 medications TEXT DEFAULT '[]', documents TEXT DEFAULT '[]',
                 open_to_trials INTEGER DEFAULT 0, created_at)

hospital_accounts(id, username, password_hash, hospital_name, location,
                  research_conditions TEXT DEFAULT '[]', created_at)

patient_trial_interests(id, patient_id, trial_id, trial_title, match_score,
                        status DEFAULT 'interested', created_at,
                        UNIQUE(patient_id, trial_id))

connections(id, patient_id, hospital_id, trial_id, trial_title,
            initiated_by DEFAULT 'patient', status DEFAULT 'pending',
            message, created_at,
            UNIQUE(patient_id, hospital_id, trial_id))
-- Post-schema index handles NULL trial_id:
-- CREATE UNIQUE INDEX idx_conn_unique ON connections(patient_id, hospital_id, COALESCE(trial_id, ''))

connection_messages(id, connection_id, sender_role TEXT,  -- 'patient' or 'hospital'
                    sender_id TEXT, body TEXT,
                    created_at, is_read INTEGER DEFAULT 0)
-- Index: idx_msgs_conn ON connection_messages(connection_id, created_at)
```

JSON fields (conditions, medications, documents, research_conditions) are stored as TEXT and parsed via `_row_to_dict()`.

### Key database.py Functions

| Function | Purpose |
|---------|---------|
| `get_open_patients_for_hospital(hid, condition_filter, include_connected)` | Returns open_to_trials=1 patients; when `include_connected=False` excludes already-connected patients |
| `get_connection_messages(connection_id)` | All messages for a connection, ASC order |
| `create_connection_message(connection_id, sender_role, sender_id, body)` | Insert message, returns dict |
| `mark_messages_read(connection_id, reader_role)` | Mark all messages from the other role as read |
| `unread_count(connection_id, reader_role)` | Count of unread messages from the other role |
| `get_hospital_inbox_threads(hospital_id)` | All threads with last_message, last_message_at, unread_count; sorted by activity |
| `get_patient_inbox_threads(patient_id)` | Same for patient side |
| `_seed_dataset_patients(c, max_patients=20)` | Seeds Synthea patients from CSV on first run |

---

## Feature Engineering (17 features in FEATURE_COLS)

| Feature | Description |
|---------|-------------|
| condition_overlap | Raw count of shared conditions |
| jaccard_similarity | overlap / union |
| overlap_ratio_trial | overlap / len(trial_conditions) |
| overlap_ratio_patient | overlap / len(patient_conditions) |
| condition_rarity_score | mean(1/log2(n_trials_per_cond+2)), normalised 0-1 |
| trial_specificity | 1 / trial_condition_count |
| condition_burden | total_patient_conds / 10 |
| active_ratio | active_conds / total_conds |
| resolved_ratio | resolved_conds / total_conds |
| age_distance | normalised distance outside age range (0 if within) |
| age_centered | position within age range (-1 to +1) |
| age_compatibility | 1.0 in range, decays over 30-year gap |
| gender_compatibility | 1.0 match/all, 0.1 mismatch |
| **geo_feasibility** | **State-level: 1.0 same state, 0.75 other US state, 0.5 no data** |
| **med_compatibility** | **Keyword overlap: patient meds vs trial drug interventions** |
| **lab_availability** | **Patient observation/lab type coverage (0-1, normalised by 20)** |
| data_completeness | fraction of key fields present |

---

## Flask API Routes

### Auth (no login required)
- `POST /auth/patient/register` β†’ {success, patient} or {error}
- `POST /auth/patient/login` β†’ {success, patient} or {error}
- `POST /auth/hospital/register` β†’ {success, hospital} or {error}
- `POST /auth/hospital/login` β†’ {success, hospital} or {error}
- `POST /auth/logout` β†’ {success}

### Patient API (requires patient session)
- `GET  /api/patient/profile` β†’ patient dict (no password_hash)
- `POST /api/patient/profile` β†’ updated patient dict; allowed fields: first_name, last_name, dob, gender, address, conditions, medications, open_to_trials
- `GET  /api/patient/matches` β†’ {results: [...trials], total}
- `GET  /api/patient/interests` β†’ {interests: [...]}
- `POST /api/patient/interest` β†’ {success}; body: {trial_id, trial_title, match_score}
- `DELETE /api/patient/interest/<trial_id>` β†’ {success}
- `GET  /api/patient/connections` β†’ {connections: [...]} joined with hospital_name
- `POST /api/patient/connect` β†’ {success, connection} or 409 if duplicate; body: {hospital_id, trial_id, trial_title, message}
- `GET  /api/patient/hospitals-for-trial?trial_id=NCT...` β†’ {hospitals: [...]} tiered matching (see below)
- `GET  /api/patient/connections/<cid>/messages` β†’ {messages: [...]}; marks hospital messages read
- `POST /api/patient/connections/<cid>/messages` β†’ {success, message}; body: {body}
- `GET  /api/patient/inbox` β†’ {threads: [...]} each with last_message, last_message_at, unread_count, hospital_name
- `GET  /api/patient/documents` β†’ {documents: [...]}
- `POST /api/patient/documents` β†’ {success, document}; multipart/form-data file upload (max 10 MB, .pdf/.docx/.doc/.txt/.png/.jpg/.jpeg)
- `DELETE /api/patient/documents/<doc_id>` β†’ {success}; removes file from disk and DB
- `GET  /api/patient/documents/<doc_id>/download` β†’ file download (as_attachment)

### Hospital API (requires hospital session)
- `GET  /api/hospital/profile` β†’ hospital dict (no password_hash)
- `POST /api/hospital/profile` β†’ updated hospital dict; allowed fields: hospital_name, location, research_conditions
- `GET  /api/hospital/patients?condition=&include_connected=` β†’ {patients: [...]} open_to_trials=1; `include_connected=true` to include already-connected patients (used by Search tab)
- `POST /api/hospital/connect` β†’ {success, connection} or 409 if duplicate; body: {patient_id, trial_id, trial_title, message}
- `GET  /api/hospital/connections` β†’ {connections: [...]} joined with patient fields
- `PUT  /api/hospital/connections/<cid>/status` β†’ {success}; body: {status: pending|accepted|rejected|completed}
- `GET  /api/hospital/connections/<cid>/messages` β†’ {messages: [...]}; marks patient messages read
- `POST /api/hospital/connections/<cid>/messages` β†’ {success, message}; body: {body}
- `GET  /api/hospital/inbox` β†’ {threads: [...]} each with last_message, last_message_at, unread_count, first_name, last_name
- `GET  /api/hospital/trials` β†’ {trials: [...]} active trials matched to hospital profile (see below)

### Shared
- `GET /api/status` β†’ {ready, stats} or {ready: false, message}
- `GET /api/conditions/autocomplete?q=...` β†’ {results: [...]}

---

## Hospital Trial Dashboard (`/api/hospital/trials`)

`pipeline.trials_for_hospital(hospital_name, location, research_conditions, top_k=20)` β€” reverse of patient matching: given a hospital's profile, find active clinical trials it is most relevant to.

| Tier | Match condition | `match_reason` field |
|------|----------------|----------------------|
| 1 | Jaccard(hospital name tokens, trial facility name tokens) β‰₯ 0.25 | "name matched to trial site" |
| 2 | Hospital state matches a US trial facility state | "in same state as trial site" |
| 3 | Hospital `research_conditions` overlaps trial conditions | "researches related conditions" |

Active-only filter (`is_active` check) applied at every tier. Returns list of dicts:
`trial_id, title, phase, status, summary, location, facility_name, n_sites, match_tier, match_reason`

---

## Tiered Hospital Matching (`/api/patient/hospitals-for-trial`)

For each trial, hospitals in the DB are scored and returned in tier order (Tier 1 first):

| Tier | Match condition | `match_reason` field | UI label |
|------|----------------|----------------------|----------|
| 1 | Jaccard(hospital name tokens, any trial facility name tokens) β‰₯ 0.25 | "verified site on this trial" | Green β€” Verified Trial Sites |
| 2 | Hospital state (from "City, ST" location) matches a trial US facility state | "in same state as a trial site" | Grey β€” Related Hospitals |
| 3 | Hospital `research_conditions` overlaps trial conditions | "researches related conditions" | Grey β€” Related Hospitals |
| 4 | Fallback β€” trial has no facility/state data at all | "" | Grey β€” Related Hospitals |

The patient connect modal groups Tier 1 hospitals under a green "VERIFIED TRIAL SITES" header and Tiers 2-4 under a grey "RELATED HOSPITALS β€” not confirmed trial sites" header.

**Pipeline lookups used:**
- `pipeline.trial_facility_tokens[trial_id]` β€” list of frozensets of significant words from facility names
- `pipeline.trial_us_states[trial_id]` β€” set of US state full names (e.g. {"Massachusetts"})
- `pipeline.trial_profiles[trial_id]["conditions"]` β€” set of condition strings

**Stopword set for facility name tokenisation** (same in pipeline.py and app.py):
hospital, medical, center, centre, clinic, university, health, care, healthcare, system, institute, foundation, research, general, regional, national, community, services, department, division, college, school, the, of, and, at, for, in, a, an, is, by

---

## Trial Match Result Fields

Each item in `/api/patient/matches` results:
- trial_id, title, phase, status, min_age, max_age, sex, enrollment, start_date
- eligibility_probability (0-100, calibrated RF probability Γ— 100)
- match_score (0-100, rule-based: overlap_ratio weighted)
- combined_score (0-100, 0.6 Γ— eligibility + 0.4 Γ— match_score)
- age_compatibility, gender_compatibility, geo_feasibility, med_compatibility (all 0-100)
- condition_rarity_score (0-1)
- overlap_conditions (list of conditions shared with patient)
- trial_conditions (all trial conditions)
- criteria (eligibility criteria text, truncated 500 chars)
- summary (brief summary, truncated 400 chars)
- **facility_name** (lead US facility name, or "" if not available)
- location (lead US facility city/state/country string)
- n_sites (total facility count for this trial)
- interest_status (null | 'interested' | 'withdrawn', from patient_trial_interests)

---

## Data Privacy Model

1. Hospitals see only patients with open_to_trials=1 (name, age, gender, conditions)
2. Full details accessible only after patient-initiated connection
3. Hospital cannot contact a patient unless patient is open to trials
4. Connection record: patient_id, hospital_id, trial_id, initiated_by, status, message
5. Hospital can also initiate connections with opt-in patients from the hospital portal

---

## XSS Prevention

All user-controlled strings use DOM API (never innerHTML for user data):
```javascript
function escH(s) {       // text content in innerHTML contexts
  const d = document.createElement('div');
  d.appendChild(document.createTextNode(String(s||'')));
  return d.innerHTML;
}
function escA(s) {       // HTML attribute values
  return String(s||'').replace(/&/g,'&amp;').replace(/"/g,'&quot;')
                      .replace(/</g,'&lt;').replace(/>/g,'&gt;');
}
```
Event listeners use addEventListener only. Tags and cards built via createElement + textContent.

---

## ML Model

- Random Forest (n_estimators=200, max_depth=12, class_weight="balanced")
- CalibratedClassifierCV (isotonic, cv=3) for probability calibration
- GroupShuffleSplit (patient-level, 80/20, no leakage) for train/test split
- GroupKFold (5-fold, patient-level) for cross-validation
- Training sample: 3000 patients Γ— 30 trials each + random negatives
- Cache: `model_cache.pkl` (delete to force retrain)

### Actual Metrics (from verified live run, 2026-04-25)

| Metric | Value |
|--------|-------|
| Accuracy | 85.15% |
| AUC-ROC | 0.5976 |
| CV AUC (5-fold) | 0.5992 Β± 0.0052 |
| F1 | 0.9168 |
| Precision | 0.8518 |
| Recall | 0.9925 |
| Brier score | 0.1263 |
| Avg precision | 0.8540 |
| Train size | 90,994 pairs |
| Test size | 22,681 pairs |
| Positive label rate | 82.1% |

> **Note on AUC:** The 82.1% positive rate in pseudo-labels (weighted 6-feature labelling threshold at 0.5) makes the classification task easy to solve trivially β€” high accuracy/recall but lower AUC. To improve AUC, the pseudo-label threshold should be raised (e.g. 0.6) or positive/negative sampling balanced more aggressively.

### Feature Importance (Random Forest, ranked)

| Rank | Feature | Importance |
|------|---------|-----------|
| 1 | age_distance | 0.2040 |
| 2 | age_compatibility | 0.1691 |
| 3 | gender_compatibility | 0.1288 |
| 4 | age_centered | 0.0791 |
| 5 | jaccard_similarity | 0.0751 |
| 6 | condition_rarity_score | 0.0621 |
| 7 | overlap_ratio_trial | 0.0519 |
| 8 | overlap_ratio_patient | 0.0512 |
| 9 | condition_overlap | 0.0396 |
| 10 | condition_burden | 0.0274 |
| 11 | resolved_ratio | 0.0263 |
| 12 | active_ratio | 0.0258 |
| 13 | lab_availability | 0.0208 |
| 14 | geo_feasibility | 0.0158 |
| 15 | med_compatibility | 0.0144 |
| 16 | trial_specificity | 0.0087 |
| 17 | data_completeness | 0.0000 |

---

## MIMIC-IV Validation (code-level only, not in UI)

- 100 demo patients, ~90% match rate after 3-tier ICD β†’ condition mapping
- Call: `pipeline.validate_mimic()` β†’ list of {subject_id, mapped_conditions, n_matches, top_match}
- 3-tier mapping: exact β†’ substring containment β†’ word-overlap β‰₯ 75%
- Not exposed via any Flask route

---

## All Bug Fixes by Session

### Session 3 Fixes (2026-04-25) β€” Two-portal foundation

#### Fix 1 β€” patient_id not passed to match_patient() (CRITICAL)
**Before:** `pipeline.match_patient(conditions, age, gender, top_k=20)`
**After:** `pipeline.match_patient(conditions, age, gender, top_k=20, patient_id=session["patient_id"], address=address)`

Without this, patient-specific medication keywords and lab scores defaulted to empty / 0.3 for all users β€” med_compatibility and lab_availability were effectively constants.

#### Fix 2 β€” Pseudo-label used only 3 features (HIGH)
**Before:** AND gate on age/gender/condition β€” geo/med/lab had near-zero training influence.
**After:** 6-feature weighted score with 15% random noise:
```python
score = (
    0.30 * float(row["age_compatibility"]    > 0.6) +
    0.15 * float(row["gender_compatibility"] > 0.5) +
    0.25 * float(row["jaccard_similarity"]   > 0.05) +
    0.10 * float(row["geo_feasibility"]) +
    0.10 * float(row["med_compatibility"]) +
    0.10 * float(row["lab_availability"])
)
base = int(score >= 0.5)
# 15% hash-deterministic noise for realism
```

#### Fix 3 β€” geo_feasibility was country-level heuristic (MEDIUM)
**Before:** Float from `trail_countries.csv` (1.0 US, 0.7 multi-national, 0.35 non-US). No patient location.
**After:** State-level matching using `trail_facilities.csv` + patient address regex:
```python
def _geo_score(patient_state_full, trial_states):
    if not trial_states: return 0.5       # no US facility data β€” neutral
    if patient_state_full in trial_states: return 1.0
    return 0.75                           # other US state
```

#### Fix 4 β€” NameError `trial_geo` in _compute_features return dict
`"geo_feasibility": float(trial_geo)` β†’ `"geo_feasibility": geo_feasibility`

#### Fix 5 β€” XSS in condition tag onclick handlers (MEDIUM)
`addConditionTag('${c}')` broke for conditions with apostrophes (e.g. "alzheimer's disease").
Fixed with DOM-based `makeTag()` using textContent + addEventListener. No inline onclick anywhere.

#### Fix 6 β€” Duplicate connection prevention (was: no guard)
- `connections` table: added `UNIQUE(patient_id, hospital_id, trial_id)` schema constraint
- `init_db()`: runs `CREATE UNIQUE INDEX IF NOT EXISTS idx_conn_unique ON connections(patient_id, hospital_id, COALESCE(trial_id, ''))` to handle NULL trial_id and backfill existing DBs
- `create_connection()`: pre-checks `connection_exists()` before insert; returns `None` on duplicate
- `/api/patient/connect` and `/api/hospital/connect`: return 409 when `create_connection()` returns None

#### Fix 7 β€” Hospital patient feed showed already-contacted patients (was: no exclusion)
`get_open_patients_for_hospital()` now uses:
```sql
WHERE open_to_trials=1
  AND id NOT IN (SELECT DISTINCT patient_id FROM connections WHERE hospital_id=?)
```

#### Fix 8 β€” Demo seed: john_doe starts with open_to_trials=1
Hospital portal was empty on a fresh database. `_seed_demo_data()` now seeds john_doe with `open_to_trials=1`.

#### Fix 9 β€” Hospital registration silently ignored research_conditions
`templates/landing.html` hospital register form now collects comma-separated research conditions and sends them as a parsed lowercase array to the backend.

#### Fix 10 β€” Trial cards only showed site count, not facility name or location
`pipeline.py match_patient()` now extracts `facility_name` from `Facility_Name` column; prefers US facilities. Patient portal detail grid shows "Lead Site" and "Location" when available.

---

### Session 4 Fixes (2026-04-25) β€” Hospital matching overhaul + profile editing

#### Fix 11 β€” Hospital suggestion logic replaced (was: research_conditions overlap only)
Complete replacement of `/api/patient/hospitals-for-trial`:

**Before:** looped all hospitals, included any whose `research_conditions` overlapped trial conditions. No tier concept, no facility data used.

**After:** 4-tier system using two new pipeline lookups:
- `pipeline.trial_facility_tokens[trial_id]` β€” built from `Facility_Name` column in trail_facilities.csv, US rows only. Each facility name tokenised by stripping stopwords + words < 3 chars.
- `pipeline.trial_us_states[trial_id]` β€” set of full US state names for the trial

Helper functions in `app.py`:
```python
def _hospital_name_tokens(name: str) -> frozenset:
    # strips stopwords, keeps words β‰₯ 3 chars
    ...

def _facility_match_score(h_tokens: frozenset, facility_token_list: list) -> float:
    # best Jaccard score against any facility in the trial
    ...
```

Each hospital gets one tier assigned and a `match_reason` + `match_tier` in the response.
Result list sorted by `match_tier` ascending (best first).

#### Fix 12 β€” Hospital portal had no profile editing
`POST /api/hospital/profile` added (was GET-only). `database.py update_hospital_profile()` added. `templates/hospital.html` now has a **My Profile** tab with editable hospital name, location, and research condition tags. On save, the navbar hospital name updates live without a page reload.

#### Fix 13 β€” Model disclaimer missing from patient trial results
`templates/patient.html` trial results section now shows an alert above results:
> "Match percentages are predictions from a model trained on synthetic patient data and rule-based labels β€” not validated clinical eligibility determinations. Always consult a healthcare provider before enrolling in any trial."

---

### Session 5 Fixes (2026-04-25) β€” Modal UX + bug fixes

#### Fix 14 β€” Patient connect modal showed all hospitals in one flat list
**Before:** All hospitals (all tiers) in a single flat list, sorted by tier, with coloured badges as the only visual distinction.

**After:** Modal renders two visually separated sections:
- **"VERIFIED TRIAL SITES"** (green `sec-head`) β€” Tier 1 hospitals only
- **"RELATED HOSPITALS β€” not confirmed trial sites"** (grey `sec-head` with inline subtitle) β€” Tiers 2, 3, 4

Both sections only render if they have entries. Click delegation on the outer `#hospitalList` wrapper still works for both sections.

#### Fix 15 β€” Close button invisible on hospital profile condition tags
`hospital.html renderProfTags()`: `btn-close-white` (white X) on `bg-info text-dark` badge (light blue background) β†’ `btn-close` (dark X). The X was invisible before.

#### Fix 16 β€” Login/register forms required mouse click, no Enter key support
`templates/landing.html`: Added `_onEnter(inputId, fn)` helper and wired Enter key on all login and register inputs (both patient and hospital portals). Works on username field too (not just password).

#### Fix 17 β€” System ready banner never auto-cleared on slow boot
`templates/patient.html`: `checkStatus()` was called once at DOMContentLoaded and never again. If the pipeline was still training when the user opened the page, the yellow banner persisted even after the pipeline finished.

**After:** `startStatusPoll()` starts a `setInterval` (5s) when the initial check finds `ready: false`. The interval clears itself once `ready: true` is received.
```javascript
checkStatus().then(() => { if (!sysReady) startStatusPoll(); });
```

---

## Live Run Verification (2026-04-25)

End-to-end test results after full retrain with no model_cache.pkl:

| Test | Result |
|------|--------|
| Landing page GET / | 200 OK |
| Patient login john_doe/pass123 | OK β€” returns patient JSON |
| Hospital login mgh/mgh123 | OK β€” returns hospital JSON |
| Pipeline ready (api/status) | ready: true |
| /api/patient/matches for john_doe | 20 results, all score fields populated |
| Top match geo score for Greece trial | 50% (no US facility β€” correct) |
| Hospital browses open_to_trials patients | 25 patients visible (5 demo + 20 Synthea) |
| Hospital condition search ?condition=hypertension | results including Synthea patients |
| Hospital β†’ patient connect (POST) | OK, status=pending |
| Patient sees hospital connection (GET) | 1 connection, hospital_name present |
| /api/conditions/autocomplete?q=hyper | ["hypertension"] |
| Duplicate connect attempt | 409 error |
| Hospital profile save | navbar name updates live |
| Patient connect modal | Two sections render correctly |
| /api/hospital/trials for mgh | Active trials with match tiers |
| /api/hospital/inbox | Threads with unread counts |
| /api/patient/inbox | Threads with hospital names |
| Patient document upload | File saved, metadata in DB |
| Inbox Synthea seeding | 25 patients total confirmed |

---

### Session 6 Fixes (2026-04-25) β€” Messaging, inbox, trial dashboard, document upload, dataset seeding

#### Fix 18 β€” Hospital trial dashboard (My Trials tab)
**Before:** Hospital portal had no way to see which clinical trials were relevant to it.

**After:** New "My Trials" nav tab in `hospital.html`. Calls `GET /api/hospital/trials` β†’ `pipeline.trials_for_hospital()`. Active-only filter at all 3 tiers. Cards show status badge, phase, match tier, facility name, summary excerpt, and a "View on ClinicalTrials.gov β†—" link.

Active-only filter: `if not self.trial_profiles[trial_id].get("is_active", False): continue` at each tier loop in `pipeline.py`.

#### Fix 19 β€” Messaging (chat in connections)
**Before:** Connections table had no messaging. Patients and hospitals could only see connection status.

**After:**
- New `connection_messages` table with `sender_role`, `sender_id`, `body`, `is_read`.
- `GET/POST /api/patient/connections/<cid>/messages` and `GET/POST /api/hospital/connections/<cid>/messages`.
- Both portals have a messages modal (`#msgModal`) opened by a Chat button in the Connections table.
- `mark_messages_read()` called on GET to auto-mark messages as read when the recipient opens the thread.

#### Fix 20 β€” Dedicated Inbox tab (both portals)
**Before:** Chat only accessible from the My Connections table row β€” no inbox overview.

**After:** New "Inbox" nav tab in both `hospital.html` and `patient.html`.
- Calls `GET /api/hospital/inbox` or `GET /api/patient/inbox`.
- Backed by `get_hospital_inbox_threads()` / `get_patient_inbox_threads()` β€” SQL subqueries aggregate last_message, last_message_at, unread_count per thread.
- Threads sorted by most recent activity (Python-side sort on `last_message_at or created_at`).
- Unread count badge on nav tab button updates when inbox loads.
- "Open" button reuses the existing `openMsgModal()` and messages modal.

#### Fix 21 β€” Document upload (patient portal)
**Before:** Patient profile had no file upload section.

**After:** "My Documents" card added to patient profile tab. 4 routes:
- `POST /api/patient/documents` β€” werkzeug `secure_filename`, 10 MB limit, allowed extensions: `.pdf/.docx/.doc/.txt/.png/.jpg/.jpeg`. Saves to `uploads/patient_docs/<patient_id>/`. Metadata stored as JSON array in `patient_accounts.documents`.
- `DELETE /api/patient/documents/<doc_id>` β€” removes file from disk and metadata from DB.
- `GET /api/patient/documents/<doc_id>/download` β€” serves file as attachment.
- `app.config["MAX_CONTENT_LENGTH"] = 10 * 1024 * 1024` enforced Flask-side.

#### Fix 22 β€” Dataset-backed patient seeding
**Before:** Only hand-made demo patients in DB (john_doe only had open_to_trials=1 initially). Hospital search returned 0 results on fresh DB.

**After:** `_seed_dataset_patients(c, max_patients=20)` in `database.py` seeds 20 real Synthea patients from CSV files on first run. Skips deceased patients (Death_Date not empty). Reads up to 8 conditions + 6 medications per patient. Username: `synthea_<first8chars_of_Patient_ID>`, password `pass123`, `open_to_trials=1`.

All 5 hand-made demo patients also set to `open_to_trials=1`. Total: 25 patients in DB.

#### Fix 23 β€” Hospital search include_connected toggle
**Before:** Hospital Search tab also excluded already-connected patients, same as Available Patients tab β€” making it useless for re-searching.

**After:** Search tab adds `include_connected=true` query param. `GET /api/hospital/patients?include_connected=true` bypasses the exclusion subquery. Available Patients tab retains strict exclusion. Toggle checkbox in Search tab UI.

#### Fix 24 β€” Template auto-reload
**Before:** `app.run(debug=False, use_reloader=False)` β€” template edits required server restart to take effect.

**After:** `app.config["TEMPLATES_AUTO_RELOAD"] = True` added after other config lines. Templates now reload on every request without enabling full debug mode or the reloader.

---

## Known Issues / Future Improvements

1. **High pseudo-label positive rate (82.1%)** β€” lowers AUC-ROC to ~0.60. Fix: raise label threshold from 0.5 to 0.6, or explicitly sample equal positive/negative pairs.

2. **data_completeness feature importance = 0** β€” nearly constant across training pairs (all synthetic patients have complete data). Consider removing from FEATURE_COLS.

3. **Fuzzy condition matching not implemented** β€” only exact condition name overlaps used (106 conditions). Substring/semantic fuzzy matching would expand coverage significantly.

4. **lab_availability coverage is low (8.7%)** β€” observations file is sparse. Consider normalising denominator to the subset that has any lab data.

5. **Hospital portal does not rank patients by match quality** β€” listed in insertion order. Could rank by condition overlap with the hospital's research_conditions.

6. **No email/notification system** β€” connection requests visible only inside the portal.

7. **No automated tests** β€” syntax checking only (`python -m py_compile`). Key flows to cover: patient registration β†’ profile update β†’ match β†’ connect; hospital registration β†’ patient browse β†’ connect β†’ status update; duplicate connection rejection.

8. **Inbox badge not auto-refreshed** β€” unread count badge only updates when the user clicks the Inbox tab. No real-time push; would require polling or WebSockets.

9. **Document access control** β€” uploaded files are served from disk by doc_id only; no additional hospital-side access to patient documents (by design β€” privacy model). Hospital sees document count in patient profile only after connection.

10. **Synthea patients have Synthea-style names** (e.g. "Geovany567 Reichert456") β€” cosmetically odd but functionally correct. No fix needed for demo.