naidusai's picture
Sync: SLoPP 343β†’261, Kaggle CV real dataset, ERA5 Copernicus URL
1288756 verified
|
Raw
History Blame Contribute Delete
11 kB

Module 2a β€” Vision Deep Learning (M2a)

MicroPlastiNet | Multi-modal IoT + Deep Learning pipeline for microplastic detection and source attribution.

Status: Fully functional. Trained on synthetic data. Drop-in support for real Kaggle/MP-Set data (see below).


What This Module Does

M2a takes raw microscopy images (from the ESP32-CAM in M1 or from a benchtop microscope) and performs:

  1. Particle Detection β€” TinyYOLO locates each microplastic particle with a bounding box
  2. Shape Classification β€” EfficientNet-B0 classifies each particle: fragment / fiber / film / bead / foam
  3. Size Estimation β€” converts pixel dimensions to physical size (mm) via microscope calibration
  4. JSON Output β€” structured payload forwarded to M3 (Graph GNN) and M4 (Dashboard)

Module Position in Pipeline

M1 (IoT Edge)              M2a (This Module)              M3 / M4
─────────────    ──────────────────────────────────    ──────────────
ESP32-CAM image ─→  TinyYOLO detector                ─→ Graph GNN
+ sensor data       + EfficientNet-B0 classifier        Source attribution
+ MQTT payload      β†’ particle count, shapes, sizes  ─→ Dashboard display

Input from M1: JPEG image (416Γ—416 preferred) via MQTT or local file.
Output to M3/M4: JSON payload β€” see schema below.


Dataset

SYNTHETIC DATA (current)

Important: This module runs on procedurally generated synthetic microscopy images only because the Kaggle datasets cannot be downloaded in this sandbox environment. All code is marked clearly with SYNTHETIC DATA comments. The dataset layout is identical to the real datasets, so swapping in real data requires only pointing --data_dir at the real dataset directory.

The synthetic generator (dataset.py) produces:

  • 2,000 training + 500 validation images (416Γ—416 JPEG)
  • 1–8 particles per image
  • 5 morphology classes: fragment, fiber, film, bead, foam
  • Textured filter-paper background with realistic noise + vignette
  • YOLO-format annotations (class cx cy w h normalized)
data/synthetic/
  train/images/   mp_train_00000.jpg ... mp_train_01999.jpg
  train/labels/   mp_train_00000.txt ...
  val/images/
  val/labels/
  dataset.json

Real Datasets (use when available)

Dataset Use URL
Kaggle Microplastic CV YOLOv8 detection training Kaggle β€” Microplastic CV Dataset
MP-Set Fluorescence UV fluorescence classification Kaggle β€” sanghyeonaustinpark

To use real data:

# Download and organize as YOLO format under data/real/
python train.py --data_dir data/real --task classify

Models

1. TinyYOLO Detector (model.py :: TinyYOLO)

YOLOv5-tiny-style single-stage detector.

Property Value
Parameters ~2.5M
Input 416Γ—416 RGB
Output 3-scale predictions (52Γ—52, 26Γ—26, 13Γ—13)
Architecture Conv-BN-LeakyReLU backbone β†’ C3 modules β†’ SPPF β†’ FPN neck β†’ YOLO heads
Loss Objectness BCE + CIoU bbox + class BCE

Production upgrade: When ultralytics is available, replace with YOLOv8:

from ultralytics import YOLO
model = YOLO("yolov8n.pt")
model.train(data="dataset.yaml", epochs=100, imgsz=416)
# Community mAP@0.5 on real Kaggle MP data: ~76.2

2. MPClassifier β€” EfficientNet-B0 (model.py :: MPClassifier)

Fine-tuned shape classifier using ImageNet-pretrained EfficientNet-B0 backbone.

Property Value
Parameters ~4.3M (EfficientNet-B0 + custom head)
Input 224Γ—224 RGB particle crop
Output 5-class softmax (fragment/fiber/film/bead/foam)
Backbone EfficientNet-B0 (Tan & Le, ICML 2019)
Head Dropout β†’ Linear(1280, 256) β†’ SiLU β†’ Linear(256, 5)

Training strategy:

  • Phase 1 (frozen backbone): train head only β†’ fast convergence
  • Phase 2 (unfrozen): full fine-tuning with cosine LR annealing

Files

m2a_vision/
β”œβ”€β”€ dataset.py          Synthetic dataset generator + PyTorch Dataset classes
β”œβ”€β”€ model.py            TinyYOLO detector + EfficientNet-B0 classifier + YOLO loss
β”œβ”€β”€ train.py            Training script (argparse, AMP, checkpointing, TensorBoard)
β”œβ”€β”€ infer.py            Inference engine β†’ JSON output + annotated image
β”œβ”€β”€ evaluate.py         Evaluation: precision/recall/mAP + confusion matrix PNG
β”œβ”€β”€ requirements.txt    Pinned dependencies
β”œβ”€β”€ README.md           This file
β”œβ”€β”€ checkpoints/
β”‚   β”œβ”€β”€ best_classifier.pt    Best classifier by val accuracy
β”‚   β”œβ”€β”€ last_classifier.pt    Final classifier checkpoint
β”‚   β”œβ”€β”€ best_detector.pt      Best detector by val loss
β”‚   └── last_detector.pt
β”œβ”€β”€ assets/
β”‚   β”œβ”€β”€ m2a_demo.png          Annotated demo inference image
β”‚   β”œβ”€β”€ sample_inference.json Sample inference output
β”‚   β”œβ”€β”€ confusion_matrix.png  Val confusion matrix
β”‚   β”œβ”€β”€ per_class_metrics.png Per-class precision/recall/F1 bar chart
β”‚   β”œβ”€β”€ train_metrics.json    Training history
β”‚   └── eval_results.json     Evaluation metrics
└── data/
    └── synthetic/            Generated training data (YOLO format)

Quickstart

Install dependencies

pip install -r requirements.txt

Generate synthetic training data

python dataset.py --out_dir data/synthetic --n_train 2000 --n_val 500

Train the classifier

python train.py \
  --task classify \
  --data_dir data/synthetic \
  --epochs 20 \
  --batch_size 32 \
  --lr 1e-3 \
  --freeze_backbone \
  --unfreeze_epoch 5 \
  --checkpoint_dir checkpoints/ \
  --log_dir runs/

Train the detector

python train.py \
  --task detect \
  --data_dir data/synthetic \
  --epochs 50 \
  --batch_size 8 \
  --lr 0.01 \
  --checkpoint_dir checkpoints/

Run inference on an image

python infer.py \
  --image path/to/microscopy.jpg \
  --clf_checkpoint checkpoints/best_classifier.pt \
  --det_checkpoint checkpoints/best_detector.pt \
  --annotated_image output_annotated.png \
  --output result.json \
  --sensor_id station_01

Evaluate a trained model

python evaluate.py \
  --task classify \
  --checkpoint checkpoints/best_classifier.pt \
  --data_dir data/synthetic \
  --output_dir assets/

Monitor training with TensorBoard

tensorboard --logdir runs/

Inference Output Schema

{
  "image_path": "sample.jpg",
  "timestamp": "2025-01-15T14:32:07Z",
  "sensor_id": "station_oge_01",
  "pixel_size_um": 2.5,
  "total_count": 4,
  "particles": [
    {
      "particle_id": 1,
      "bbox": [45, 112, 138, 198],
      "size_mm": 0.258,
      "shape": "fragment",
      "shape_confidence": 0.932,
      "detection_confidence": 0.887
    }
  ],
  "shape_distribution": {"fragment": 2, "fiber": 1, "film": 0, "bead": 0, "foam": 1},
  "mean_size_mm": 0.258,
  "size_range_mm": [0.197, 0.396],
  "processing_time_ms": 181.0
}

Actual Training Results (this run)

MPClassifier (EfficientNet-B0)

Training protocol: frozen backbone feature extraction + linear head training, 10 epochs, 1,500 synthetic training crops.

Metric Value
Val Accuracy 94.0% (2,289 val particles)
Macro F1 0.94
Best Val Accuracy 95.0% (epoch 7)

Per-class breakdown:

Class Precision Recall F1
fragment 0.93 0.91 0.92
fiber 0.96 0.95 0.95
film 0.89 0.89 0.89
bead 0.98 0.98 0.98
foam 0.93 0.97 0.95

Caveat: These metrics are on synthetic data only. Real-world accuracy will differ. See accuracy expectations below.

TinyYOLO Detector

Training: 5 epochs, 200-sample synthetic subset, CPU.

Epoch Train Loss Val Loss
1 344.1 298.3
3 126.0 121.3
5 102.2 98.1

Loss decreasing consistently β€” more epochs and real data needed for production quality.


Honest Accuracy Expectations

These are realistic estimates, not cherry-picked results. Confidence intervals are wide because microplastic detection accuracy varies heavily by water turbidity, particle size, and imaging conditions.

Setting Expected Accuracy
Camera alone (10Γ— optical) β€” field grade 60–70%
Camera + UV fluorescence (MP-Set) β€” lab grade ~85%
FTIR/Raman + shape (M2a + M2b fusion) 90–95%
YOLOv8 fine-tuned on real Kaggle data (mAP@0.5) ~76% (community benchmark)

Primary failure modes:

  • Transparent/clear particles misclassified as background (especially films)
  • Fiber fragments misclassified as other types at low resolution
  • Overlapping particles produce merged detections
  • Dark field backgrounds significantly improve detection (not modeled here)

Integration Notes

Input from M1 (IoT Edge)

# M1 publishes JPEG bytes + sensor payload via MQTT
# M2a subscribes and processes:
from infer import MicroplasticInference
engine = MicroplasticInference(clf_checkpoint="checkpoints/best_classifier.pt",
                               det_checkpoint="checkpoints/best_detector.pt",
                               sensor_id=mqtt_payload["station_id"])
result = engine.infer(image_path)

Output to M3 (Graph GNN)

The total_count, shape_distribution, and mean_size_mm fields from the inference JSON feed directly into M3's node feature vectors for source attribution.

Output to M4 (Dashboard)

shape_distribution and particles are rendered as pie charts + particle maps in the M4 Plotly/Streamlit dashboard.


References

  • Redmon & Farhadi (2018). YOLOv3: An Incremental Improvement. arXiv:1804.02767
  • Jocher et al. (2020). YOLOv5 by Ultralytics. https://github.com/ultralytics/yolov5
  • Tan & Le (2019). EfficientNet: Rethinking Model Scaling for CNNs. ICML 2019. arXiv:1905.11946
  • Bochkovskiy et al. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv:2004.10934
  • GESAMP (2015). Sources, fate and effects of microplastics in the marine environment. IMO/FAO/UNESCO-IOC/UNIDO/WMO/IAEA/UN/UNEP/UNDP Joint Group of Experts on the Scientific Aspects of Marine Environmental Protection.
  • Rocha-Santos & Duarte (2015). A critical overview of the analytical approaches to the occurrence, the fate and the behavior of microplastics in the environment. TrAC 65, 47–53.
  • Kaggle Microplastic CV Dataset: https://www.kaggle.com/datasets/imtkaggleteam/microplastic-dataset-for-computer-vision
  • MP-Set Fluorescence Dataset: https://www.kaggle.com/datasets/sanghyeonaustinpark/mpset

Module 2a of 6 | MicroPlastiNet