File size: 12,670 Bytes
79b0bef
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184d5a3
79b0bef
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
# Clinical NLP Pipeline β€” VSCode Setup & Run Guide

Everything you need to go from a fresh clone to a running
Streamlit demo in one sitting.

---

## Prerequisites

| Tool | Version | Install |
|------|---------|---------|
| Python | 3.10+ | python.org |
| Git | any | git-scm.com |
| VSCode | any | code.visualstudio.com |
| Tesseract (optional) | 5.x | see Phase 0 |

---

## Phase 0 β€” Clone and virtual environment

```bash
git clone https://github.com/ayodeji07/clinical-nlp-pipeline.git
cd clinical-nlp-pipeline

python -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

pip install --upgrade pip
pip install -r requirements-dev.txt
```

Copy the example env file:

```bash
cp .env.example .env
```

Open `.env` β€” for local development the defaults are fine.
`DATABASE_URL` will default to SQLite automatically.

---

## Phase 1 β€” Install NLP models

The scispaCy model is not on PyPI β€” install it directly:

```bash
pip install scispacy==0.5.5 --no-deps
pip install conllu pysbd "nmslib-metabrainz==2.1.3"
pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_lg-0.5.4.tar.gz --no-deps
```

> **Note**: `--no-deps` is required on spaCy β‰₯ 3.8 because scispaCy pins `spacy<3.8.0`.
> The model loads and runs correctly despite the version mismatch.

Verify:

```bash
python -c "import spacy; nlp = spacy.load('en_core_sci_lg'); print('NER model OK')"
```

---

## Phase 2 β€” Download MTSamples

1. Go to https://www.kaggle.com/datasets/tboyle10/medicaltranscriptions
2. Download `mtsamples.csv`
3. Save it to `data/raw/mtsamples.csv`

Also download the ICD-10 reference file (included in the repo under `data/raw/`).
If it is missing, download from:
https://www.cms.gov/medicare/coding-billing/icd-10-codes

Save as `data/raw/icd10_codes.csv` with columns: `code`, `description`.

---

## Phase 3 β€” Run the ETL pipeline

Dry run first to validate data without writing to the database:

```bash
python -m src.etl.pipeline --dry-run
```

Full run:

```bash
python -m src.etl.pipeline
```

Expected output:
```
Stage 1: Extract β€” loaded 4,999 raw notes
Stage 2: Transform β€” 4,721 notes ready
Stage 3: Load β€” inserted 4,721 records
```

The ETL pipeline only loads notes β€” it doesn't run NER, ICD-10 mapping,
or classification against them. Populate those with the batch scripts:

```bash
python scripts/run_ner_batch.py                            # extract entities from stored notes
python scripts/run_icd10_batch.py                          # map DISEASE/SYMPTOM entities to ICD-10 codes
python scripts/train_severity_classifier.py --n-seeds 4    # fine-tune the severity classifier
```

All three are idempotent and resumable, so interrupting and re-running
is safe. Pass `--limit N` to any of them for a quick subset run instead
of processing the full dataset. The dashboard's Explorer and Model
Metrics pages need this step done first β€” without it there's nothing
to show beyond raw note counts.

`--n-seeds N` on the classifier script trains N different random seeds
and keeps only the checkpoint with the best critical-class F1, then
records it in the `model_runs` table (what `GET /model/metrics` and
the Model Metrics dashboard page read from). Fine-tuning a small
classification head on a small dataset is sensitive to random
init β€” an unseeded single run isn't reliably comparable across
retrains, so this matters more than it might look like it should.
Omit the flag (or use `--n-seeds 1`) for a single quick run.

If your Supabase database was created before this feature was added,
`model_runs` is missing the columns this needs (`create_all_tables()`
only creates missing tables, not missing columns on existing ones) β€”
run the migration in `sql/schema.sql` (the commented `ALTER TABLE`
statements near the bottom of the `model_runs` block) once via the
Supabase SQL Editor, or execute them directly:

```bash
python -c "
from sqlalchemy import text
from src.db.connection import get_session
stmts = [
    'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS test_accuracy REAL',
    'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS test_f1 REAL',
    'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS per_class JSON',
    'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS confusion_matrix JSON',
    'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS history JSON',
]
with get_session() as session:
    for stmt in stmts:
        session.execute(text(stmt))
"
```

---

## Phase 4 β€” Run the notebooks

Open VSCode, install the Jupyter extension, then open notebooks in order:

```
notebooks/00_data_exploration.ipynb   ← understand the dataset
notebooks/01_ner_walkthrough.ipynb    ← try the NER pipeline
notebooks/02_icd_mapping.ipynb        ← try ICD-10 mapping
notebooks/03_classification.ipynb     ← train the classifier (~20 min CPU)
notebooks/04_visualisation.ipynb      ← build the charts
```

Select kernel: Python (``.venv``)

> **GPU tip**: For faster training, open notebook 03 in Google Colab.
> Upload the notebook, mount your Drive, and run β€” takes ~5 minutes on T4.

---

## Phase 5 β€” Start the API

```bash
uvicorn src.api.main:app --reload --port 8000
```

Open http://localhost:8000/docs to see the Swagger UI.

Test the health endpoint:

```bash
curl http://localhost:8000/health
# β†’ {"status":"ok","version":"1.0.0","database":"connected"}
```

Test the analyse endpoint:

```bash
curl -X POST http://localhost:8000/notes/analyse \
  -H "Content-Type: application/json" \
  -d '{"text": "Patient has hypertension and takes metformin.", "include_icd10": true}'
```

---

## Phase 6 β€” Start the Streamlit dashboard

In a second terminal (keep the API running in the first):

```bash
streamlit run dashboard/app.py
```

Open http://localhost:8501

You should see:
- πŸ”¬ **Demo** page β€” paste a note and click Analyse
- πŸ“Š **Explorer** β€” entity frequency + co-occurrence graph
- 🧠 **Metrics** β€” classifier performance (after training notebook 03)

---

## Phase 7 β€” Run the tests

```bash
pytest
```

Run with coverage:

```bash
pytest --cov=src --cov-report=term-missing
```

Expected: 65 tests pass, ~0 failures.

---

## Phase 8 β€” Deploy to Streamlit Cloud

### Step 1: Create a Supabase database

1. Go to https://supabase.com β†’ New project
2. Choose a region close to your users
3. Copy the connection string from Settings β†’ Database β€” but use the
   **Pooler** connection string (Transaction or Session mode), not the
   direct one. It looks like:
   `postgresql://postgres.<ref>:<password>@aws-<region>.pooler.supabase.com:5432/postgres`

   **Do not use** the direct-connection host
   (`db.<ref>.supabase.co`) for a deployed app β€” it resolves to an
   IPv6-only address, and several free-tier hosts (Render's included)
   don't support IPv6 egress. You'll get a "Network is unreachable"
   error at startup that looks like a code bug but is actually just
   the wrong host. The pooler host is IPv4-compatible and works
   everywhere.

### Step 2: Deploy the API

The API needs to be publicly accessible for Streamlit Cloud to reach it.

- **Hugging Face Spaces** (recommended for this project) β€”
  https://huggingface.co/new-space, SDK: **Docker**. Push this repo to
  the Space's git remote and it builds from the existing `Dockerfile`.
  The free CPU tier has enough RAM (historically ~16GB) to run the full
  hybrid NER pipeline + ICD-10 embeddings + classifier together without
  issue β€” confirmed working in production for this project. Sleeps
  after 48h of inactivity, not 15 minutes, so it stays warm for normal
  demo traffic.

- **Railway** β€” https://railway.app (free tier)
  ```bash
  railway login
  railway up
  ```
  Copy the public URL Railway gives you.

- **Render** β€” https://render.com (free tier) β€” **be aware its free
  tier caps out at 512Mi RAM**, which is not enough to run this
  project's full hybrid NER pipeline (`en_ner_bc5cdr_md` +
  `en_core_sci_lg`) together with the classifier β€” a real
  `/notes/analyse` request will get OOM-killed (502 Bad Gateway),
  even though `/health` and DB-backed endpoints work fine. If you use
  Render anyway:
  - Set `WARM_UP_MODELS=false` so the app can at least boot instead of
    crash-looping on eager model load at startup.
  - Set `NER_MODEL=en_ner_bc5cdr_md` (single model instead of
    `hybrid`) to meaningfully cut memory, at the cost of losing
    `PROCEDURE`/`ANATOMY`/most `SYMPTOM` entity coverage (bc5cdr alone
    still detects `DISEASE`/`MEDICATION` β€” the more important two).
  - Or upgrade to a paid instance size.

Set `DATABASE_URL` (the pooler string from Step 1) as an environment
variable on whichever platform you choose.

### Step 3: Deploy the Streamlit app

1. Go to https://share.streamlit.io β†’ New app
2. Connect your GitHub repo
3. Set **Main file path**: `dashboard/app.py`
4. Click **Advanced settings** β†’ **Secrets** and add:

```toml
API_BASE_URL = "https://<your-username>-<space-name>.hf.space"
DATABASE_URL = "postgresql://postgres.<ref>:<password>@aws-<region>.pooler.supabase.com:5432/postgres"
```

(Swap `API_BASE_URL` for your Railway/Render URL if you went that
route instead.)

5. Click **Deploy**

Your app will be live at `https://<your-app>.streamlit.app`


---

## Phase 9 β€” Docker (optional)

Run the full stack in containers without installing anything locally
except Docker Desktop.

```bash
# Build and start API + dashboard
docker-compose up --build

# API: http://localhost:8000/docs
# Dashboard: http://localhost:8501
```

Stop everything:

```bash
docker-compose down
```

The `data/` directory is mounted as a volume so your database and
processed files persist between container restarts.

---

## Supabase β€” applying the schema

When you create a new Supabase project the database is empty.
SQLAlchemy will create the tables automatically on first API startup,
but you can also apply the schema manually for review or migration:

1. Open your Supabase project dashboard
2. Click **SQL Editor** in the left sidebar
3. Click **New query**
4. Open `sql/schema.sql` from this repo and paste the contents
5. Click **Run** (Ctrl+Enter)

You should see all four tables appear in the **Table Editor**:
`clinical_notes`, `entities`, `icd10_mappings`, `model_runs`.

> **Note**: The schema uses `AUTOINCREMENT` which is SQLite syntax.
> For Supabase (PostgreSQL), replace `INTEGER PRIMARY KEY AUTOINCREMENT`
> with `SERIAL PRIMARY KEY` in the SQL editor. SQLAlchemy handles
> this automatically when creating tables via `create_all_tables()`.

---

## Setting up Streamlit secrets locally

Streamlit Cloud reads secrets from `.streamlit/secrets.toml`.
For local development, create this file (it is in `.gitignore`):

```bash
mkdir -p .streamlit
cat > .streamlit/secrets.toml << 'EOF'
API_BASE_URL = "http://localhost:8000"
# DATABASE_URL = "postgresql://..."  # only needed for cloud
EOF
```

The dashboard will read `API_BASE_URL` from here automatically.
---

## Troubleshooting

**`OSError: en_core_sci_lg not found`**
```bash
pip install scispacy==0.5.5 --no-deps
pip install conllu pysbd "nmslib-metabrainz==2.1.3"
pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_lg-0.5.4.tar.gz --no-deps
```

**`FileNotFoundError: mtsamples.csv not found`**
Download from Kaggle and save to `data/raw/mtsamples.csv`.

**`ModuleNotFoundError: No module named 'src'`**
Run commands from the repo root, or add it to PYTHONPATH:
```bash
export PYTHONPATH=$(pwd)
```

**`RuntimeError: Model not loaded`**
Run notebook 03 to fine-tune and save the classifier before
calling `clf.predict()`.

**Dashboard shows "API offline"**
Ensure `uvicorn src.api.main:app --reload` is running in a
separate terminal and `API_BASE_URL` in `.env` points to it.

**Supabase connection timeout / "Network is unreachable" on a deployed host**
Make sure you're using the *pooler* connection string, not the direct
one β€” the direct host (`db.<ref>.supabase.co`) is IPv6-only and
unreachable from several free-tier hosts (confirmed on Render's free
tier). Use the pooler host instead, with `?sslmode=require` appended:
```
postgresql://postgres.<ref>:<pass>@aws-<region>.pooler.supabase.com:5432/postgres?sslmode=require
```

---

## Project structure (quick reference)

```
src/utils/      config, logger, text cleaning
src/etl/        extract, transform, load, pipeline
src/nlp/        ner, icd_mapper, classifier, cooccurrence
src/db/         connection, models, repository
src/api/        FastAPI app + routes
dashboard/      Streamlit app + pages
notebooks/      00–04 walkthrough notebooks
tests/          pytest test suite (65 tests)
```