morinousagi's picture
add ui snapshot
a49f4b5 verified
|
Raw
History Blame Contribute Delete
6.64 kB
---
title: Pytorch ViT - SEM Images Defect Detection
emoji: πŸ”
colorFrom: indigo
colorTo: purple
sdk: docker
app_file: app.py
pinned: false
---
## SEM Images Defect Detection using ViT
The primary purpose of this project was to learn to apply **ViT**, and experiment any approach not done in classroom. This project involved prompt engineering and was completed with vibe-coding assistance from **Gemini Flash 3**.
- **Background**: Semiconductor defect images are scarce & proprietary, and are typically high-res Scanning Electron Microscopy (SEM) images. Similar to images used in this project. Pre-trained models are typically trained with images not as high-res as SEM images.
- Workarounds for data scarcity (not applied in this project): Data augmentation or using models to generate synthetic defect images, etc.
- **Model training approach**: One-class learning, train on 'normal' data only - [HF app demo](https://huggingface.co/spaces/morinousagi/pytorch-vit-defect-detect) | [UI snapshot](app_ui.png)
(Alternate training approach using sample size data - [GitHub repo](https://huggingface.co/spaces/morinousagi/pytorch-vit-defect-detect/blob/main/README.md)
### Data source
- Citation: Cheng, Deruo, 2021, "MIIC (Microscopic Images of Integrated Circuits) Dataset for Anomaly Detection and Image Inpainting", https://doi.org/10.21979/N9/WBLTFI, DR-NTU (Data), V1.
The data accompanies the publication: "Joint Anomaly Detection and Inpainting for Microscopy Images via Deep Self-Supervised Learning", IEEE International Conference on Image Processing (ICIP), 2021.
- File: <Anomaly_train.rar> over 20K normal images and 116 anomaly images -- as it is computationally intensive, only a small subset of this data was used in this project for practice and demo purpose only.
### Dataset
- Convert 512x512 grayscale (1 channel) to RGB (3 channels) as expected by the pre-trained ViT
- Scaling (handles by ToTensor)
Normalization is not required as there is image reconstruction within Autoencoder. Decoder with Sigmoid() will output pixels in [0, 1] -> target to be in the same range as output
```
/data (local)
/train
/defective (0 images)
/normal (3,000 images)
/test
/defective (116 images)
/normal (100 images)
```
### Model
Unsupervised (reconstructive) anomaly detection
- Autoencoders - a specialized, self-supervised type (no labels) of encoder-decoder model focused on reconstructing their own input to learn efficient data representations (input = output).
- init
- Load pre-trained ViT as encoder: `google/vit-base-patch16-224` (expects 224p resolution by default)
- Custom convolutional decoder to reconstruct the image - mapping ViT features back to image space
- forward pass
- Requires position interpolation as input images are maintained at 512p resolution to prevent loss of info (i.e. tiny defects)
### Training & Evaluation
- Optimizer: AdamW (improves upon Adam by decoupling weight decay from the gradient update process, for better regularization and generalization)
- Learning rate: 1e-4
- Loss: MSE
- Evaluation: Peak score logic (1st attempt used average scoring did not yield good results)
### Pipeline
Run ```python pipeline.py```
- Set global variables
- Run dataloaders, training and evaluation
- Calculate best threshold using F1-score
- Save model file
### Results
Trained with 3000 'normal' images (topographic and metal interconnects) for 15 epochs only - non-concluding
```
--- Starting Training (Normal Images Only) ---
Epoch 1: 0%| | 0/375 [00:00<?, ?it/s]
Epoch 1: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 1 complete. Avg Loss: 0.007519
Epoch 2: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:11<00:00, 1.95s/it]
Epoch 2 complete. Avg Loss: 0.000961
Epoch 3: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 3 complete. Avg Loss: 0.000684
Epoch 4: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 4 complete. Avg Loss: 0.000546
Epoch 5: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 5 complete. Avg Loss: 0.000473
Epoch 6: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:13<00:00, 1.96s/it]
Epoch 6 complete. Avg Loss: 0.000410
Epoch 7: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:13<00:00, 1.96s/it]
Epoch 7 complete. Avg Loss: 0.000370
Epoch 8: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 8 complete. Avg Loss: 0.000345
Epoch 9: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 9 complete. Avg Loss: 0.000305
Epoch 10: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:13<00:00, 1.96s/it]
Epoch 10 complete. Avg Loss: 0.000278
Epoch 11: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 11 complete. Avg Loss: 0.000272
Epoch 12: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 12 complete. Avg Loss: 0.000245
Epoch 13: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 13 complete. Avg Loss: 0.000235
Epoch 14: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:12<00:00, 1.95s/it]
Epoch 14 complete. Avg Loss: 0.000223
Epoch 15: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 375/375 [12:13<00:00, 1.96s/it]
Epoch 15 complete. Avg Loss: 0.000212
--- Evaluating & Finding Threshold ---
Computing anomaly scores (Peak Score Method)...
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 27/27 [00:17<00:00, 1.54it/s]
Optimal Anomaly Threshold: 0.0016485939268022776
Max F1-Score: 0.8433
βœ… Pipeline Complete. Download 'model.safetensors' and 'config.json' from the file pane.
```
![image](avg_loss.png)
- Re-calibrated Threshold to 0.0025
- Observation: Model had not learned certain metal layer features of integrated circuits such as via/contact, 'shadow' effect around elevated structures, metal OPC edges.
### Repo
```
sem-vit-anomaly/ # Model repo
β”œβ”€β”€ config.json
└── model.safetensors
pytorch-vit-defect-detect/ # This repo
β”œβ”€β”€ src/
β”‚ └── __init__.py # Explicitly mark src as a contained module
β”‚ └── dataset.py # Image transformation and dataloaders
β”‚ └── model.py # SEMViTAutoencoder class
β”‚ └── train_eval.py # Training and evaluation functions
β”‚ └── pipeline.py # Run pipeline
β”œβ”€β”€ app.py # App deployment
β”œβ”€β”€ Dockerfile # App deployment
└── requirements.txt # App deployment
```
### Ref
- Gemini Flash 3
- https://huggingface.co/google/vit-base-patch16-224