--- title: Heart Attack Risk Predictor emoji: ๐Ÿซ€ colorFrom: blue colorTo: purple sdk: docker app_file: app.py pinned: false ---

๐Ÿซ€ Heart Attack Risk Predictor โ€” Multimodal

An AI-powered clinical decision-support tool that estimates a patient's heart-attack risk from tabular patient data, an ECG image, or both โ€” using an ensemble-router architecture with a tabular model, an ECG image model, and late-fusion of their scores.

๐Ÿ”ด TRY THE LIVE DEMO ON HUGGING FACE SPACES ๐Ÿ”ด

โš ๏ธ Educational / portfolio demo โ€” not a medical device. Do not use for real clinical decisions.

--- ## ๐Ÿ“‘ Table of Contents - [What's New in v2](#-whats-new-in-v2) - [Key Features](#-key-features) - [How It Works โ€” Architecture](#-how-it-works--architecture) - [The Two Models](#-the-two-models) - [Tech Stack](#-tech-stack) - [Project Structure](#-project-structure) - [Getting Started](#-getting-started) - [Usage & API](#-usage--api) - [Model Performance](#-model-performance) - [Testing](#-testing) - [Limitations & Honesty](#-limitations--honesty) - [Team Members](#-team-members) --- ## ๐Ÿ†• What's New in v2 Version 1 was a single-model biomarker classifier (8 vitals โ†’ Random Forest). A leakage test showed its ~98% accuracy was largely a biomarker-threshold rule (accuracy fell to ~62% without Troponin & CK-MB). Version 2 re-architects the project into a **multimodal ensemble router**: - **Two models** instead of one โ€” a **tabular** model and an **ECG image** model. - **A router** that picks the model(s) based on what the user submits, and **averages** their scores when both are provided. - **Honest evaluation** โ€” proper metrics (ROC-AUC for the imbalanced tabular task, per-class metrics for the ECG task) and a clear statement of limitations. > The original v1 biomarker research is preserved in `research_and_experiments/`. --- ## โœจ Key Features - **Multimodal input** โ€” enter patient data, upload an ECG image, or do both. - **Ensemble router** โ€” one `POST /predict` endpoint routes to the right model(s): tabular โ†’ Model A, ECG โ†’ Model B, both โ†’ averaged score. - **Missing-data friendly** โ€” blank tabular fields are filled automatically by **K-Nearest-Neighbours imputation**, so a partial form still works. - **Server-side validation** โ€” out-of-range or non-numeric fields are rejected with a clear message (HTTP 422). - **ECG confidence check** โ€” low-confidence ECG predictions are flagged ("may not be a clear 12-lead ECG"). - **Transparent results** โ€” the UI shows each model's score and, in "both" mode, the combined average, so nothing is a black box. - **Modern UI** โ€” dark glassmorphism theme, drag-and-drop ECG upload, color-coded risk badges (๐Ÿ”ด High / ๐ŸŸก Moderate / ๐ŸŸข Low). --- ## ๐Ÿง  How It Works โ€” Architecture ``` โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ POST /predict (multipart/form-data) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” tabular only โ”€โ”คโ†’ Model A (Framingham: KNN-impute โ†’ scale โ†’ RandomForest) โ†’ p_a โ”€โ” โ”‚ โ”‚ โ”œโ”€ both โ†’ average โ†’ p โ†’ band (Low/Mod/High) ECG only โ”€โ”€โ”€โ”€โ”€โ”คโ†’ Model B (ResNet18 transfer learning on ECG images) โ†’ p_b โ”€โ”˜ โ”‚ โ”‚ โ”‚ neither โ”€โ”€โ”€โ”€โ”€โ”€โ”คโ†’ HTTP 400 โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ ``` Each model outputs a scalar **risk probability** `p โˆˆ [0, 1]`. That value is mapped to a band โ€” **Low** (`< 0.33`), **Moderate** (`0.33โ€“0.66`), **High** (`> 0.66`) โ€” and, when both models run, the two scores are combined by an equal-weight average. --- ## ๐Ÿค– The Two Models ### Model A โ€” Tabular (Framingham 10-year CHD) - **Data:** Framingham Heart Study (4,240 patients, 15 clinical features). - **Target:** `TenYearCHD` โ€” probability of coronary heart disease within 10 years. - **Pipeline:** `StandardScaler โ†’ KNNImputer(k=5) โ†’ classifier`, saved as one artifact. - **Model selection:** RandomForest vs XGBoost by 5-fold CV ROC-AUC โ€” **RandomForest won (0.69 vs 0.67)**. - **Handles imbalance** (~15% positive) with class weighting; the app uses the continuous probability, not a hard 0.5 cutoff. ### Model B โ€” ECG image (ResNet18) - **Data:** ECG Images Dataset of Cardiac Patients (928 images, 4 classes). - **Method:** **transfer learning** โ€” a ResNet18 pretrained on ImageNet, with a new 4-class head; train the head, then fine-tune the last block. - **Classes โ†’ risk weight:** Normal `0.0`, abnormal heartbeat `0.5`, post-MI history `0.8`, myocardial infarction `1.0`. The 4-class softmax is collapsed to one risk score via these weights. --- ## ๐Ÿ› ๏ธ Tech Stack | Layer | Technology | Purpose | |---|---|---| | **Backend** | [FastAPI](https://fastapi.tiangolo.com/) + [Uvicorn](https://www.uvicorn.org/) | Async web framework + ASGI server; multipart `/predict` router | | **Tabular ML** | [scikit-learn](https://scikit-learn.org/) (`RandomForest`, `KNNImputer`, `StandardScaler`), [XGBoost](https://xgboost.readthedocs.io/) | Model A pipeline + model comparison | | **Image ML** | [PyTorch](https://pytorch.org/) + [torchvision](https://pytorch.org/vision/) (ResNet18) | Model B transfer learning | | **Images / Uploads** | [Pillow](https://python-pillow.org/), `python-multipart` | ECG image decoding + file uploads | | **Data** | [pandas](https://pandas.pydata.org/), [NumPy](https://numpy.org/) | Data handling | | **Frontend** | HTML5, CSS3, Vanilla JS | Single-page glassmorphism UI with ECG drag-and-drop | | **Deployment** | Docker โ†’ Hugging Face Spaces | Containerized serving on port 7860 | --- ## ๐Ÿ“ Project Structure ``` heart-attack-risk-predictor/ โ”‚ โ”œโ”€โ”€ app.py # FastAPI app + ensemble router (multipart /predict) โ”œโ”€โ”€ requirements.txt # Production dependencies โ”œโ”€โ”€ Dockerfile # Container build (copies app, inference/, models/, static/) โ”‚ โ”œโ”€โ”€ inference/ # Serving-time prediction package โ”‚ โ”œโ”€โ”€ fusion.py # Risk bands + late-fusion (combine) โ”‚ โ”œโ”€โ”€ framingham.py # Model A inference (with KNN imputation) โ”‚ โ”œโ”€โ”€ ecg.py # Model B inference (softmax โ†’ risk scalar) โ”‚ โ””โ”€โ”€ validation.py # Server-side field validation โ”‚ โ”œโ”€โ”€ models/ # Trained artifacts (Git LFS) โ”‚ โ”œโ”€โ”€ framingham_pipeline.joblib # Model A bundle โ”‚ โ”œโ”€โ”€ ecg_resnet.pt # Model B weights โ”‚ โ””โ”€โ”€ ecg_classes.json # ECG class list + risk weights + preprocessing โ”‚ โ”œโ”€โ”€ train_framingham.py # Trains Model A โ”œโ”€โ”€ train_ecg.py # Trains Model B โ”‚ โ”œโ”€โ”€ static/index.html # Frontend (form + ECG upload + per-branch results) โ”‚ โ”œโ”€โ”€ DOCUMENTATION.md # Full technical documentation โ”œโ”€โ”€ DEFENSE_GUIDE.md # Beginner-friendly project defense guide โ”‚ โ””โ”€โ”€ research_and_experiments/ # v1 biomarker research (notebook, dataset, old model) ``` --- ## ๐Ÿš€ Getting Started ### Prerequisites - Python **3.11+** (3.12 recommended) ### Install ```bash git clone https://github.com/Shahd1Sayed/heart-attack-risk-predictor.git cd heart-attack-risk-predictor pip install -r requirements.txt ``` ### Data (only needed to (re)train โ€” download from Kaggle) ``` data/framingham.csv # "Framingham Heart Study dataset" data/ecg_data//*.jpg # "ECG Images Dataset of Cardiac Patients" ``` ### Train (produces the files in models/) ```bash python train_framingham.py python train_ecg.py ``` ### Run ```bash uvicorn app:app --host 127.0.0.1 --port 8000 # open http://127.0.0.1:8000 ``` The app boots even before models are trained; a branch needing an untrained model returns HTTP 503 with a hint. --- ## โ–ถ๏ธ Usage & API `POST /predict` accepts **multipart/form-data** with optional tabular fields and an optional `ecg` image file. ```bash # tabular only curl -F age=61 -F male=1 -F sysBP=150 -F totChol=240 http://127.0.0.1:8000/predict # ECG only curl -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predict # both (multimodal) curl -F age=61 -F sysBP=150 -F ecg=@some_ecg.jpg http://127.0.0.1:8000/predict ``` **Tabular fields:** `male, age, education, currentSmoker, cigsPerDay, BPMeds, prevalentStroke, prevalentHyp, diabetes, totChol, sysBP, diaBP, BMI, heartRate, glucose` (all optional; blanks are KNN-imputed). **Example response (multimodal):** ```json { "mode": "multimodal", "risk_level": "High", "p_risk": 0.7306, "branches": { "tabular": { "p_risk": 0.4825, "band": "Moderate", "detail": {"CHD": 0.4825, "No CHD": 0.5175}, "imputed_fields": ["glucose"] }, "ecg": { "p_risk": 0.9787, "band": "High", "ecg_class": "myocardial_infarction_ecg_images", "confidence": 0.9587, "low_confidence": false, "warning": null } } } ``` **Status codes:** `200` success ยท `400` no input / unreadable image ยท `422` invalid tabular field ยท `503` model not trained yet. --- ## ๐Ÿ“Š Model Performance | Model | Metric | Value | |---|---|---| | **Model A โ€” Framingham** | CV ROC-AUC | **0.69** | | | Hold-out ROC-AUC | 0.64 | | | Accuracy | 0.85 *(โ‰ˆ all-negative baseline โ€” see note)* | | **Model B โ€” ECG ResNet18** | Test accuracy | **0.90** | | | Macro F1 | 0.90 | | | MI precision | 1.00 | > โš ๏ธ **Metric honesty:** Framingham is imbalanced (~15% positive), so 85% accuracy is > essentially the "always predict no-CHD" baseline โ€” which is why we lead with **ROC-AUC**. > Predicting a decade ahead from basic clinical features is genuinely hard, so ~0.64โ€“0.69 > AUC is expected. The ECG numbers are strong for a small dataset but optimistic vs. > other acquisition setups (the images are photos of printed ECGs). --- ## ๐Ÿงช Testing The project ships with a test harness (router paths, response invariants, adversarial inputs, determinism, an ECG serving sweep, input validation, ECG confidence, and concurrency). **Result: 62/62 checks pass.** See `DOCUMENTATION.md` for details. --- ## โš–๏ธ Limitations & Honesty - **The two models predict different things** โ€” Model A estimates 10-year *prognosis*; Model B classifies the *current* ECG. They are trained on **different, unpaired populations**, so the combined score is a **transparent heuristic demonstrating the architecture, not a validated clinical measure**. - **The fusion weights are hand-set** (equal average) because no paired dataset exists to learn/validate them. A learned fusion on paired data (e.g. PTB-XL) is future work. - **ECG images are photos of printouts** โ€” a small, imbalanced dataset; the CNN may not generalize to other setups. - **No out-of-distribution rejection** โ€” a non-ECG image is still classified (now flagged low-confidence, but not refused). - **Not for clinical use.** --- ## ๐Ÿ‘ฅ Team Members | Name | Role | GitHub | |---|---|---| | **Shahd Sayed** | Machine Learning Engineer | [@Shahd1Sayed](https://github.com/Shahd1Sayed) | | **Shahd Mohammed** | Full-Stack Developer | | ---