File size: 14,492 Bytes
24b19ad
 
 
 
 
 
 
 
f9f794f
 
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
 
f9f794f
 
 
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
f9f794f
c6c61ce
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
f9f794f
c6c61ce
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
f9f794f
 
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
f9f794f
 
 
 
 
c6c61ce
f9f794f
 
 
c6c61ce
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
 
 
 
 
c6c61ce
f9f794f
 
 
c6c61ce
f9f794f
 
 
 
c6c61ce
 
 
 
 
 
 
 
 
f9f794f
 
 
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
 
f9f794f
 
c6c61ce
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
 
 
 
 
 
 
 
f9f794f
 
c6c61ce
f9f794f
 
c6c61ce
 
 
 
 
 
f9f794f
 
 
 
 
 
c6c61ce
 
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
 
c6c61ce
 
 
 
 
 
 
 
f9f794f
 
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
f9f794f
c6c61ce
f9f794f
c6c61ce
f9f794f
c6c61ce
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f9f794f
 
 
 
c6c61ce
 
 
 
 
f9f794f
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
---
title: NeuroVision API
emoji: ๐Ÿง 
colorFrom: blue
colorTo: indigo
sdk: docker
app_port: 7860
---
<div align="center">

# ๐Ÿง  NeuroVision โ€” Medical Visual Question Answering

**An end-to-end VQA system for medical brain imaging, powered by BLIP, LangGraph, and FastAPI.**

[![Framework](https://img.shields.io/badge/Model-BLIP-0052CC?style=for-the-badge&logo=salesforce&logoColor=white)](https://huggingface.co/Salesforce/blip-vqa-base)
[![PyTorch](https://img.shields.io/badge/Built_with-PyTorch-EE4C2C?style=for-the-badge&logo=pytorch&logoColor=white)](https://pytorch.org/)
[![FastAPI](https://img.shields.io/badge/API-FastAPI-009688?style=for-the-badge&logo=fastapi&logoColor=white)](https://fastapi.tiangolo.com/)
[![LangGraph](https://img.shields.io/badge/Agent-LangGraph-764ABC?style=for-the-badge&logo=langchain&logoColor=white)](https://github.com/langchain-ai/langgraph)
[![DVC](https://img.shields.io/badge/Pipeline-DVC-13ADC7?style=for-the-badge&logo=dvc&logoColor=white)](https://dvc.org/)
[![Docker](https://img.shields.io/badge/Deploy-Docker-2496ED?style=for-the-badge&logo=docker&logoColor=white)](https://www.docker.com/)
[![MLflow](https://img.shields.io/badge/Tracking-MLflow-0194E2?style=for-the-badge&logo=mlflow&logoColor=white)](https://mlflow.org/)

<img src="https://github.com/user-attachments/assets/2d6bf915-cecf-4482-bde1-859eaf5fa399" width="100%">

</div>

---

## ๐Ÿ” Overview

NeuroVision is a Visual Question Answering (VQA) system designed for **medical brain imaging** (CT & MRI scans). It fine-tunes the [BLIP](https://huggingface.co/Salesforce/blip-vqa-base) model on the [VQA-RAD](https://huggingface.co/datasets/flaviagiammarino/vqa-rad) dataset and serves predictions through a high-performance **FastAPI** backend, complemented by a **LangGraph-powered AI agent** that can search PubMed and the web for medical context.

### โœจ Key Features

| Feature | Description |
|---|---|
| ๐Ÿ–ผ๏ธ **Image-based VQA** | Upload a brain scan and ask natural-language questions about it |
| ๐Ÿค– **Medical AI Agent** | LangGraph ReAct agent with PubMed + Web search tools |
| โšก **FastAPI Backend** | Async API with automatic OpenAPI docs at `/docs` |
| ๐Ÿ“Š **MLflow Tracking** | Full experiment tracking with metrics, params, and model artifacts |
| ๐Ÿ”„ **DVC Pipeline** | Reproducible data processing โ†’ training โ†’ evaluation pipeline |
| ๐Ÿณ **Docker Ready** | One-command containerized deployment |
| ๐ŸŽจ **Streamlit UI** | Interactive frontend for visual question answering |

---

## ๐Ÿ“‹ Table of Contents

- [Overview](#-overview)
- [Architecture](#-architecture)
- [Quick Start](#-quick-start)
- [Installation](#-installation)
- [Configuration](#๏ธ-configuration)
- [API Reference](#-api-reference)
- [LangGraph Agent](#-langgraph-agent)
- [Training Pipeline](#-training-pipeline)
- [Docker Deployment](#-docker-deployment)
- [Project Structure](#-project-structure)
- [Results & Future Work](#-results--future-work)

---

## ๐Ÿ— Architecture

```mermaid
graph LR
    subgraph Frontend
        A[Streamlit UI]
    end
    subgraph Backend
        B[FastAPI Server]
        C[BLIP VQA Model]
        D[LangGraph Agent]
    end
    subgraph External
        E[PubMed API]
        F[SerpAPI Web Search]
        G[Groq LLM - Gemma2]
    end

    A -- Image + Question --> B
    A -- Chat Query --> B
    B -- /predict/ --> C
    B -- /chat/ --> D
    D --> G
    D --> E
    D --> F
```

---

## ๐Ÿš€ Quick Start

Get the entire project running with **two commands**:

```bash
# 1. Clone the repository
git clone https://github.com/Aryan-coder-student/NeuroVision-BHPC-VQA.git
cd NeuroVision-BHPC-VQA

# 2. Run the automated setup script โ€” this does EVERYTHING for you
bash setup.sh
```

That's it. The `setup.sh` script handles the complete environment bootstrap:

| Step | What it does |
|---|---|
| 1๏ธโƒฃ | Installs [uv](https://astral.sh/uv) (ultra-fast Python package manager) if not already present |
| 2๏ธโƒฃ | Creates a `.venv` virtual environment |
| 3๏ธโƒฃ | Activates the virtual environment (cross-platform: Windows & Linux/Mac) |
| 4๏ธโƒฃ | Installs all dependencies from `requirements.txt` via `uv pip install` |
| 5๏ธโƒฃ | Downloads the [VQA-RAD dataset](https://huggingface.co/datasets/flaviagiammarino/vqa-rad) from Hugging Face into `data/bronze/` |

### After Setup

```bash
# Add your API keys (required for the medical chatbot)
echo "SERPAPI_API_KEY=your_key_here" > Deployment/.env
echo "GROQ_API_KEY=your_key_here" >> Deployment/.env

# Activate the virtual environment (if not already active)
# Windows
.venv\Scripts\activate
# Linux / macOS
source .venv/bin/activate

# Launch the API server
python Deployment/app.py
```

The API will be live at **`http://localhost:5000`** with interactive Swagger docs at **`http://localhost:5000/docs`**.

---

## ๐Ÿ“ฆ Installation

### Prerequisites

| Requirement | Version |
|---|---|
| Python | 3.10+ |
| CUDA (optional) | 11.8+ (for GPU acceleration) |
| Git | 2.30+ |
| [uv](https://astral.sh/uv) | Latest (auto-installed by `setup.sh`) |

### Manual Setup

```bash
# 1. Create and activate virtual environment
python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux / macOS
source .venv/bin/activate

# 2. Install dependencies
pip install -r requirements.txt

# 3. Download the VQA-RAD dataset
python -c "
from datasets import load_dataset
dataset = load_dataset('flaviagiammarino/vqa-rad')
dataset.save_to_disk('data/bronze/')
print('Dataset downloaded to data/bronze/')
"
```

### Environment Variables

Create a `Deployment/.env` file:

```env
SERPAPI_API_KEY=your_serpapi_key        # Required for Medical Web Search tool
GROQ_API_KEY=your_groq_api_key         # Required for the LLM (Gemma2-9b-it)
```

| Variable | Purpose | Get it from |
|---|---|---|
| `SERPAPI_API_KEY` | Powers the Medical Web Search tool | [serpapi.com](https://serpapi.com/) |
| `GROQ_API_KEY` | Powers the Gemma2 LLM via Groq | [console.groq.com](https://console.groq.com/) |

---

## โš™๏ธ Configuration

All project configuration is centralized in two YAML files:

**`config.yaml`** โ€” Model paths and data locations:
```yaml
finetune_model:
  best: models/best-saved-model
  last: models/last-saved-model
  orignal_model_id: Salesforce/blip-vqa-base
data_location:
  data: data/bronze/flaviagiammarino___vqa-rad
  train_processed_data: data/silver/train_dataset.pkl
  test_processed_data: data/silver/test_dataset.pkl
result: results
```

**`param.yaml`** โ€” Training hyperparameters:
```yaml
params:
  batch_size: 8
  num_epochs: 50
  learning_rate: 5e-5
  weight_decay: 1e-4
  gradient_accumulation_steps: 4
  patience: 10
```

---

## ๐Ÿ“ก API Reference

The FastAPI server exposes two endpoints. Full interactive documentation is auto-generated at **`/docs`** (Swagger UI) and **`/redoc`** (ReDoc).

### `POST /predict/` โ€” Image Question Answering

Upload a medical image and ask a question about it.

**Request** (multipart form-data):
```bash
curl -X POST "http://localhost:5000/predict/" \
  -F "file=@brain_scan.jpg" \
  -F "question=Is there a tumor visible?"
```

**Response:**
```json
{
  "answer": "yes"
}
```

### `POST /chat/` โ€” Medical AI Chatbot

Ask medical questions powered by the LangGraph agent.

**Request** (JSON):
```bash
curl -X POST "http://localhost:5000/chat/" \
  -H "Content-Type: application/json" \
  -d '{"query": "What are the latest treatment options for glioblastoma?"}'
```

**Response:**
```json
{
  "response": "Glioblastoma treatment typically involves a multimodal approach including surgical resection, radiation therapy (usually 60 Gy in 30 fractions), and concurrent temozolomide chemotherapy..."
}
```

### Python Client Example

```python
import requests

# Image VQA
files = {"file": open("brain_scan.jpg", "rb")}
data = {"question": "What abnormality is present?"}
response = requests.post("http://localhost:5000/predict/", files=files, data=data)
print(response.json())

# Medical Chat
payload = {"query": "Explain the differences between CT and MRI for brain imaging"}
response = requests.post("http://localhost:5000/chat/", json=payload)
print(response.json())
```

---

## ๐Ÿค– LangGraph Agent

The medical chatbot uses a modern **LangGraph ReAct agent** architecture โ€” a stateful, graph-based agent that reasons step-by-step and calls tools as needed.

### Architecture

| Component | Technology | Purpose |
|---|---|---|
| **LLM** | Groq Gemma2-9b-it | Fast inference for reasoning and response generation |
| **Agent Framework** | LangGraph `create_react_agent` | Graph-based ReAct loop with tool calling |
| **Memory** | `MemorySaver` checkpointer | Persists conversation history across requests |
| **Web Search** | SerpAPI | Real-time medical web search |
| **Literature Search** | PubMed API | Peer-reviewed research paper retrieval |

### How It Works

```
User Query โ†’ LangGraph Agent โ†’ Reason โ†’ Select Tool(s) โ†’ Execute โ†’ Synthesize โ†’ Response
                  โ†‘                                                      |
                  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Memory (MemorySaver) โ†โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
```

1. The agent receives the user query along with conversation history
2. The LLM reasons about which tools to invoke (or responds directly)
3. Tools are called (PubMed, web search) and results are collected
4. The LLM synthesizes a comprehensive answer from tool outputs
5. Conversation state is persisted via the `MemorySaver` checkpointer

---

## ๐Ÿ”ฌ Training Pipeline

The full training pipeline is managed by **DVC** for reproducibility and **MLflow** for experiment tracking.

### Pipeline Stages

```
data/bronze/ โ”€โ”€โ†’ preprocess โ”€โ”€โ†’ data/silver/ โ”€โ”€โ†’ train โ”€โ”€โ†’ models/ โ”€โ”€โ†’ evaluate โ”€โ”€โ†’ results/
```

| Stage | Script | Input | Output |
|---|---|---|---|
| **Preprocess** | `src/preprocess_data.py` | Raw VQA-RAD dataset | Tokenized pickle files |
| **Train** | `src/train.py` | Processed data + params | Fine-tuned BLIP model |
| **Evaluate** | `src/evaluate.py` | Trained model + test data | BLEU scores & metrics |

### Training Features

- ๐Ÿ”ฅ **Mixed Precision Training** (FP16) for memory efficiency
- ๐Ÿ“ˆ **Gradient Accumulation** (4 steps) to simulate larger batch sizes
- ๐Ÿ›‘ **Early Stopping** with configurable patience
- ๐Ÿ“Š **MLflow Tracking** for all metrics, params, and model artifacts
- ๐Ÿ”„ **Learning Rate Warmup** with linear decay scheduling
- โœ‚๏ธ **Gradient Clipping** (max_norm=1.0) for training stability

### Run the Pipeline

```bash
# Execute all stages
dvc repro

# Run individual stages
dvc repro preprocess
dvc repro train
dvc repro evaluate

# Push data to remote storage
dvc push

# Pull data from remote storage
dvc pull
```

### View Experiment Tracking

```bash
mlflow ui
# Open http://localhost:5000 to view experiments
```

---

## ๐Ÿณ Docker Deployment

### Build and Run

```bash
# Build the image
docker build -t neurovision-vqa .

# Run the container
docker run -p 5000:5000 \
  -e SERPAPI_API_KEY=your_key \
  -e GROQ_API_KEY=your_key \
  neurovision-vqa
```

### Dockerfile

```dockerfile
FROM python:3.10
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
EXPOSE 5000
CMD ["python", "Deployment/app.py"]
```

### Run with Streamlit UI

To run both the API and the Streamlit frontend simultaneously:

```bash
# Terminal 1 โ€” API Server
python Deployment/app.py

# Terminal 2 โ€” Streamlit UI
streamlit run Deployment/streamlit/main.py --server.port=8501
```

The Streamlit UI will be available at **`http://localhost:8501`**.

---

## ๐Ÿ“ Project Structure

```
NeuroVision-BHPC-VQA/
โ”œโ”€โ”€ .dvc/                         # DVC configuration
โ”œโ”€โ”€ .agents/                      # Agent workflows and skills
โ”‚   โ””โ”€โ”€ workflows/
โ”‚       โ””โ”€โ”€ git-push.md           # Git workflow for this project
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ bronze/                   # Raw VQA-RAD dataset (DVC-tracked)
โ”‚   โ””โ”€โ”€ silver/                   # Preprocessed tokenized data
โ”œโ”€โ”€ Deployment/
โ”‚   โ”œโ”€โ”€ app.py                    # FastAPI application entry point
โ”‚   โ”œโ”€โ”€ test_api.py               # API integration tests
โ”‚   โ””โ”€โ”€ streamlit/
โ”‚       โ””โ”€โ”€ main.py               # Streamlit frontend
โ”œโ”€โ”€ models/
โ”‚   โ”œโ”€โ”€ best-saved-model/         # Best checkpoint (by BLEU score)
โ”‚   โ””โ”€โ”€ last-saved-model/         # Latest checkpoint
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ model.py                  # BLIP model loading & configuration
โ”‚   โ”œโ”€โ”€ preprocess_data.py        # Dataset preprocessing pipeline
โ”‚   โ”œโ”€โ”€ train.py                  # Training loop with MLflow tracking
โ”‚   โ”œโ”€โ”€ evaluate.py               # BLEU score evaluation
โ”‚   โ””โ”€โ”€ trl_rlhf_train.py        # Experimental RLHF training script
โ”œโ”€โ”€ results/                      # Evaluation outputs (JSON)
โ”œโ”€โ”€ mlruns/                       # MLflow experiment data
โ”œโ”€โ”€ config.yaml                   # Model & data path configuration
โ”œโ”€โ”€ param.yaml                    # Training hyperparameters
โ”œโ”€โ”€ dvc.yaml                      # DVC pipeline definition
โ”œโ”€โ”€ dvc.lock                      # DVC pipeline lock file
โ”œโ”€โ”€ Dockerfile                    # Container build configuration
โ”œโ”€โ”€ requirements.txt              # Python dependencies
โ”œโ”€โ”€ setup.sh                      # Automated environment setup script
โ””โ”€โ”€ README.md
```

---

## ๐Ÿ“Š Results & Future Work

### Current Capabilities

- โœ… Fine-tuned BLIP model on VQA-RAD for medical image Q&A
- โœ… BLEU score evaluation with early stopping on best checkpoint
- โœ… Full experiment reproducibility via DVC + MLflow
- โœ… Production-ready async API with FastAPI
- โœ… Conversational medical AI agent with persistent memory

### ๐Ÿ”ฎ Roadmap

- [ ] Add QLoRA 4-bit quantization for efficient deployment on consumer hardware
- [ ] RLHF fine-tuning for improved answer quality
- [ ] Multi-modal RAG with medical image retrieval
- [ ] Expand to chest X-ray and pathology datasets
- [ ] Add image segmentation overlays for explainable predictions
- [ ] Multilingual VQA support
- [ ] Deployment to Hugging Face Spaces

---

## ๐Ÿ“„ License

This project is for educational and research purposes.

---

<div align="center">

**NeuroVision** โ€” Medical Visual Question Answering

Built with โค๏ธ using BLIP ยท FastAPI ยท LangGraph ยท PyTorch

</div>