--- title: NephroScreen emoji: 🩺 colorFrom: green colorTo: blue sdk: docker app_port: 8000 pinned: false license: mit --- # NephroScreen — Early Chronic Kidney Disease Screening **An end-to-end applied machine-learning case study.** It estimates chronic kidney disease (CKD) risk from routine clinical lab values — a **leakage-safe scikit-learn pipeline**, a **FastAPI** service, and a **clean web interface**, deployable on a free tier. 📄 **[Read the case study →](CASE_STUDY.md)** · ⚕️ *Educational project — not a medical device.* Trained on a small public dataset (400 patients) and *not* validated for clinical use. --- ## Why this project is interesting (the honest version) On the UCI CKD dataset the models reach near-perfect scores — Random Forest hits **100% test accuracy**. On most portfolios that number would be presented as the headline. Here it is treated as a **teaching moment**, which is the more defensible engineering story: - The dataset is small (400 rows) and the classes are **highly separable** on a handful of biomarkers (hemoglobin, PCV, serum creatinine). 100% is a *property of the data*, not evidence of a deployable clinical model. - The interesting engineering is therefore **not** the accuracy — it's doing the workflow *correctly*: no data leakage, honest validation, a screening-aware decision threshold, interpretability, and a reproducible serving path. ## Highlights | Area | What's here | |---|---| | **Leakage-safe pipeline** | Custom `CKDPreprocessor` (imputation, encoding, feature engineering) fit on the **training fold only**, wrapped in a scikit-learn `Pipeline` so CV and serving never see test statistics. | | **Screening-aware threshold** | Decision threshold tuned to keep **recall ≥ 0.99** — a missed CKD case is the costly error in screening. | | **Interpretability** | Per-prediction, rule-based clinical indicators in the API; global + local **SHAP** analysis in the notebook. | | **Serving** | Trained pipeline persisted with `joblib`; **FastAPI** loads it once and serves JSON predictions. | | **Frontend** | Dependency-free HTML/CSS/JS single page; imputes missing fields so a partial lab panel still works. | | **Reproducible** | Pinned dependencies; fixed seeds; automated `pytest` suite. | ## Model performance Held-out test set (80 patients), 10-fold stratified CV, `random_state=42`: | Model | Accuracy | Precision | Recall | F1 | ROC-AUC | 10-Fold CV | |---|---|---|---|---|---|---| | Logistic Regression | 98.75% | 100% | 98% | 98.99% | 1.0000 | 99.38% ± 1.25% | | **Random Forest** (served) | **100%** | 100% | 100% | 100% | 1.0000 | 98.75% ± 2.86% | | Decision Tree | 97.50% | 98% | 98% | 98% | 0.9893 | 96.56% ± 2.19% | Top Random Forest predictors: hemoglobin, packed cell volume, serum creatinine, `kidney_stress_index` (engineered), red blood cell count. ## Interpretability & maturity Full write-up in the **[case study](CASE_STUDY.md)**. In short: | SHAP — why the model decides | Calibration | Screening threshold | |---|---|---| | ![SHAP](assets/shap_summary.png) | ![Calibration](assets/calibration.png) | ![Threshold](assets/threshold.png) | - **SHAP** attributions are clinically coherent (low hemoglobin/PCV, high creatinine, hypertension/diabetes drive CKD predictions). - **Calibration** curve checks the probabilities are trustworthy, not just the labels. - **Recall-tuned threshold (0.62)** — for screening, a missed case costs more than a false alarm, so the served threshold keeps recall ≥ 0.99. ## Architecture ``` Browser (frontend/) ──POST /api/predict──► FastAPI (api/) ──► joblib pipeline (models/) │ CKDPreprocessor → RandomForest ``` ``` src/nephroscreen/ config · preprocessing · train (the ML package) api/ main (FastAPI) · schemas (pydantic) (the service) frontend/ index.html · style.css · app.js (the UI) models/ ckd_pipeline.joblib · metrics.json (trained artifacts) notebooks/ analysis.ipynb (SHAP + calibration) (the analysis) tests/ test_api.py ``` ## Quickstart ```bash python -m venv .venv && source .venv/Scripts/activate # Windows Git Bash pip install -r requirements.txt && pip install -e . python -m nephroscreen.train # trains + saves models/ckd_pipeline.joblib uvicorn api.main:app --reload # http://127.0.0.1:8000 pytest # run the test suite ``` Open for the UI, or for the interactive API docs. ### API ```bash curl -X POST http://127.0.0.1:8000/api/predict \ -H "Content-Type: application/json" \ -d '{"hemo": 9.5, "sc": 3.2, "al": 3, "htn": "yes", "dm": "yes"}' ``` ```json { "prediction": "CKD", "probability": 0.98, "risk_band": "High", "threshold": 0.62, "key_indicators": [ ... ] } ``` ## Deployment (free tier) - **One box (recommended):** deploy the whole app on **Render** (`render.yaml` included) or any Docker host (`Dockerfile` included). FastAPI serves both the API and the static frontend. - **Split hosting:** put `frontend/` on **Vercel/Netlify** (static) and the API on **Render**; set `window.NEPHRO_API_BASE` in the frontend to the API URL. - Free instances sleep when idle; expect a ~30–50s cold start on first request. ## Limitations & future work - Single small, single-source dataset → near-perfect scores do **not** transfer to real clinical populations. - Future: external validation on an independent cohort; a **minimal low-cost feature subset** for resource-limited screening; fairness analysis across age and sex; hyperparameter tuning; probability calibration in the serving path. ## Credits Solo project by **Sadiq Mansoor** — see [CREDITS.md](CREDITS.md). An independent rebuild and extension of an earlier university coursework analysis. Dataset: UCI ML Repository, Chronic Kidney Disease (ID 336). Licensed under [MIT](LICENSE).