Spaces:
Running
Clinical NLP Pipeline β VSCode Setup & Run Guide
Everything you need to go from a fresh clone to a running Streamlit demo in one sitting.
Prerequisites
| Tool | Version | Install |
|---|---|---|
| Python | 3.10+ | python.org |
| Git | any | git-scm.com |
| VSCode | any | code.visualstudio.com |
| Tesseract (optional) | 5.x | see Phase 0 |
Phase 0 β Clone and virtual environment
git clone https://github.com/ayodeji07/clinical-nlp-pipeline.git
cd clinical-nlp-pipeline
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install --upgrade pip
pip install -r requirements-dev.txt
Copy the example env file:
cp .env.example .env
Open .env β for local development the defaults are fine.
DATABASE_URL will default to SQLite automatically.
Phase 1 β Install NLP models
The scispaCy model is not on PyPI β install it directly:
pip install scispacy==0.5.5 --no-deps
pip install conllu pysbd "nmslib-metabrainz==2.1.3"
pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_lg-0.5.4.tar.gz --no-deps
Note:
--no-depsis required on spaCy β₯ 3.8 because scispaCy pinsspacy<3.8.0. The model loads and runs correctly despite the version mismatch.
Verify:
python -c "import spacy; nlp = spacy.load('en_core_sci_lg'); print('NER model OK')"
Phase 2 β Download MTSamples
- Go to https://www.kaggle.com/datasets/tboyle10/medicaltranscriptions
- Download
mtsamples.csv - Save it to
data/raw/mtsamples.csv
Also download the ICD-10 reference file (included in the repo under data/raw/).
If it is missing, download from:
https://www.cms.gov/medicare/coding-billing/icd-10-codes
Save as data/raw/icd10_codes.csv with columns: code, description.
Phase 3 β Run the ETL pipeline
Dry run first to validate data without writing to the database:
python -m src.etl.pipeline --dry-run
Full run:
python -m src.etl.pipeline
Expected output:
Stage 1: Extract β loaded 4,999 raw notes
Stage 2: Transform β 4,721 notes ready
Stage 3: Load β inserted 4,721 records
The ETL pipeline only loads notes β it doesn't run NER, ICD-10 mapping, or classification against them. Populate those with the batch scripts:
python scripts/run_ner_batch.py # extract entities from stored notes
python scripts/run_icd10_batch.py # map DISEASE/SYMPTOM entities to ICD-10 codes
python scripts/train_severity_classifier.py --n-seeds 4 # fine-tune the severity classifier
All three are idempotent and resumable, so interrupting and re-running
is safe. Pass --limit N to any of them for a quick subset run instead
of processing the full dataset. The dashboard's Explorer and Model
Metrics pages need this step done first β without it there's nothing
to show beyond raw note counts.
--n-seeds N on the classifier script trains N different random seeds
and keeps only the checkpoint with the best critical-class F1, then
records it in the model_runs table (what GET /model/metrics and
the Model Metrics dashboard page read from). Fine-tuning a small
classification head on a small dataset is sensitive to random
init β an unseeded single run isn't reliably comparable across
retrains, so this matters more than it might look like it should.
Omit the flag (or use --n-seeds 1) for a single quick run.
If your Supabase database was created before this feature was added,
model_runs is missing the columns this needs (create_all_tables()
only creates missing tables, not missing columns on existing ones) β
run the migration in sql/schema.sql (the commented ALTER TABLE
statements near the bottom of the model_runs block) once via the
Supabase SQL Editor, or execute them directly:
python -c "
from sqlalchemy import text
from src.db.connection import get_session
stmts = [
'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS test_accuracy REAL',
'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS test_f1 REAL',
'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS per_class JSON',
'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS confusion_matrix JSON',
'ALTER TABLE model_runs ADD COLUMN IF NOT EXISTS history JSON',
]
with get_session() as session:
for stmt in stmts:
session.execute(text(stmt))
"
Phase 4 β Run the notebooks
Open VSCode, install the Jupyter extension, then open notebooks in order:
notebooks/00_data_exploration.ipynb β understand the dataset
notebooks/01_ner_walkthrough.ipynb β try the NER pipeline
notebooks/02_icd_mapping.ipynb β try ICD-10 mapping
notebooks/03_classification.ipynb β train the classifier (~20 min CPU)
notebooks/04_visualisation.ipynb β build the charts
Select kernel: Python (.venv)
GPU tip: For faster training, open notebook 03 in Google Colab. Upload the notebook, mount your Drive, and run β takes ~5 minutes on T4.
Phase 5 β Start the API
uvicorn src.api.main:app --reload --port 8000
Open http://localhost:8000/docs to see the Swagger UI.
Test the health endpoint:
curl http://localhost:8000/health
# β {"status":"ok","version":"1.0.0","database":"connected"}
Test the analyse endpoint:
curl -X POST http://localhost:8000/notes/analyse \
-H "Content-Type: application/json" \
-d '{"text": "Patient has hypertension and takes metformin.", "include_icd10": true}'
Phase 6 β Start the Streamlit dashboard
In a second terminal (keep the API running in the first):
streamlit run dashboard/app.py
You should see:
- π¬ Demo page β paste a note and click Analyse
- π Explorer β entity frequency + co-occurrence graph
- π§ Metrics β classifier performance (after training notebook 03)
Phase 7 β Run the tests
pytest
Run with coverage:
pytest --cov=src --cov-report=term-missing
Expected: 65 tests pass, ~0 failures.
Phase 8 β Deploy to Streamlit Cloud
Step 1: Create a Supabase database
Go to https://supabase.com β New project
Choose a region close to your users
Copy the connection string from Settings β Database β but use the Pooler connection string (Transaction or Session mode), not the direct one. It looks like:
postgresql://postgres.<ref>:<password>@aws-<region>.pooler.supabase.com:5432/postgresDo not use the direct-connection host (
db.<ref>.supabase.co) for a deployed app β it resolves to an IPv6-only address, and several free-tier hosts (Render's included) don't support IPv6 egress. You'll get a "Network is unreachable" error at startup that looks like a code bug but is actually just the wrong host. The pooler host is IPv4-compatible and works everywhere.
Step 2: Deploy the API
The API needs to be publicly accessible for Streamlit Cloud to reach it.
Hugging Face Spaces (recommended for this project) β https://huggingface.co/new-space, SDK: Docker. Push this repo to the Space's git remote and it builds from the existing
Dockerfile. The free CPU tier has enough RAM (historically ~16GB) to run the full hybrid NER pipeline + ICD-10 embeddings + classifier together without issue β confirmed working in production for this project. Sleeps after 48h of inactivity, not 15 minutes, so it stays warm for normal demo traffic.Railway β https://railway.app (free tier)
railway login railway upCopy the public URL Railway gives you.
Render β https://render.com (free tier) β be aware its free tier caps out at 512Mi RAM, which is not enough to run this project's full hybrid NER pipeline (
en_ner_bc5cdr_md+en_core_sci_lg) together with the classifier β a real/notes/analyserequest will get OOM-killed (502 Bad Gateway), even though/healthand DB-backed endpoints work fine. If you use Render anyway:- Set
WARM_UP_MODELS=falseso the app can at least boot instead of crash-looping on eager model load at startup. - Set
NER_MODEL=en_ner_bc5cdr_md(single model instead ofhybrid) to meaningfully cut memory, at the cost of losingPROCEDURE/ANATOMY/mostSYMPTOMentity coverage (bc5cdr alone still detectsDISEASE/MEDICATIONβ the more important two). - Or upgrade to a paid instance size.
- Set
Set DATABASE_URL (the pooler string from Step 1) as an environment
variable on whichever platform you choose.
Step 3: Deploy the Streamlit app
- Go to https://share.streamlit.io β New app
- Connect your GitHub repo
- Set Main file path:
dashboard/app.py - Click Advanced settings β Secrets and add:
API_BASE_URL = "https://<your-username>-<space-name>.hf.space"
DATABASE_URL = "postgresql://postgres.<ref>:<password>@aws-<region>.pooler.supabase.com:5432/postgres"
(Swap API_BASE_URL for your Railway/Render URL if you went that
route instead.)
- Click Deploy
Your app will be live at https://<your-app>.streamlit.app
Phase 9 β Docker (optional)
Run the full stack in containers without installing anything locally except Docker Desktop.
# Build and start API + dashboard
docker-compose up --build
# API: http://localhost:8000/docs
# Dashboard: http://localhost:8501
Stop everything:
docker-compose down
The data/ directory is mounted as a volume so your database and
processed files persist between container restarts.
Supabase β applying the schema
When you create a new Supabase project the database is empty. SQLAlchemy will create the tables automatically on first API startup, but you can also apply the schema manually for review or migration:
- Open your Supabase project dashboard
- Click SQL Editor in the left sidebar
- Click New query
- Open
sql/schema.sqlfrom this repo and paste the contents - Click Run (Ctrl+Enter)
You should see all four tables appear in the Table Editor:
clinical_notes, entities, icd10_mappings, model_runs.
Note: The schema uses
AUTOINCREMENTwhich is SQLite syntax. For Supabase (PostgreSQL), replaceINTEGER PRIMARY KEY AUTOINCREMENTwithSERIAL PRIMARY KEYin the SQL editor. SQLAlchemy handles this automatically when creating tables viacreate_all_tables().
Setting up Streamlit secrets locally
Streamlit Cloud reads secrets from .streamlit/secrets.toml.
For local development, create this file (it is in .gitignore):
mkdir -p .streamlit
cat > .streamlit/secrets.toml << 'EOF'
API_BASE_URL = "http://localhost:8000"
# DATABASE_URL = "postgresql://..." # only needed for cloud
EOF
The dashboard will read API_BASE_URL from here automatically.
Troubleshooting
OSError: en_core_sci_lg not found
pip install scispacy==0.5.5 --no-deps
pip install conllu pysbd "nmslib-metabrainz==2.1.3"
pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_lg-0.5.4.tar.gz --no-deps
FileNotFoundError: mtsamples.csv not found
Download from Kaggle and save to data/raw/mtsamples.csv.
ModuleNotFoundError: No module named 'src'
Run commands from the repo root, or add it to PYTHONPATH:
export PYTHONPATH=$(pwd)
RuntimeError: Model not loaded
Run notebook 03 to fine-tune and save the classifier before
calling clf.predict().
Dashboard shows "API offline"
Ensure uvicorn src.api.main:app --reload is running in a
separate terminal and API_BASE_URL in .env points to it.
Supabase connection timeout / "Network is unreachable" on a deployed host
Make sure you're using the pooler connection string, not the direct
one β the direct host (db.<ref>.supabase.co) is IPv6-only and
unreachable from several free-tier hosts (confirmed on Render's free
tier). Use the pooler host instead, with ?sslmode=require appended:
postgresql://postgres.<ref>:<pass>@aws-<region>.pooler.supabase.com:5432/postgres?sslmode=require
Project structure (quick reference)
src/utils/ config, logger, text cleaning
src/etl/ extract, transform, load, pipeline
src/nlp/ ner, icd_mapper, classifier, cooccurrence
src/db/ connection, models, repository
src/api/ FastAPI app + routes
dashboard/ Streamlit app + pages
notebooks/ 00β04 walkthrough notebooks
tests/ pytest test suite (65 tests)