Gaze-LIPE / docs /DOMAIN_ADAPTATION_PLAN.md
thanhhuyvan's picture
Initial release of LIPE V2 GOLD
a10ba7f
|
Raw
History Blame Contribute Delete
3.45 kB
# Domain Adaptation Plan: LIPE V2 (MPIIGaze → Gaze360)
## 1. Executive Summary
The goal of this plan is to bridge the "Domain Gap" between controlled laboratory data (MPIIGaze) and in-the-wild data (Gaze360). Our baseline zero-shot evaluation showed an MAE of **24.3°** on Gaze360. We aim to reduce this to **< 7.0°** while maintaining a CPU-optimized inference speed of > 400 FPS.
## 2. Technical Strategies & Predicted Impact
| Phase | Strategy | Technical Mechanism | Predicted MAE |
| :--- | :--- | :--- | :---: |
| 0 | **Baseline** | Zero-shot from MPIIGaze | 24.3° |
| 1 | **Domain Randomization** | Resolution Dropout + Noise Augmentation | 18.5° |
| 2 | **Structural Adaptation** | AdaLN + Feature-Level Alignment (DANN) | 12.5° |
| 3 | **Expert Guidance** | Multi-Domain Knowledge Distillation | 8.5° |
| 4 | **Geometry Constraints** | Self-Supervised Consistency Loss | **< 7.0°** |
## 3. Detailed Implementation Roadmap
### Phase 1: Data Hardening (Augmentation & Preprocessing)
* **Resolution Dropout:** Randomly downsample 16x16 patches to 8x8 during training to simulate camera distance. Use `cv2.INTER_CUBIC` to prevent upsampling artifacts.
* **Edge-Preserving Denoising:** Integrate **Bilateral Filter** (`cv2.bilateralFilter`) to smooth homogeneous regions and reduce grain without blurring the critical iris boundaries.
* **Tuned CLAHE Preprocessing:** Apply CLAHE with a conservative **clipLimit (1.0 - 1.2)** to enhance contrast in dark regions without amplifying noise into artifacts.
* **Interpolation Strategy:** Force `INTER_CUBIC` for all resizing to preserve the "Gaze Gradient" in low-resolution conditions.
* **Geometric Jitter:** Add Gaussian noise to MediaPipe landmarks (±2px) to simulate tracking instability.
* **Extreme Pose Augmentation:** Use existing rotation logic to synthetically create head poses beyond ±20°.
### Phase 2: Feature Alignment (DANN & AdaLN)
* **DANN (Domain-Adversarial Neural Network):**
* Implement a **Gradient Reversal Layer (GRL)**.
* Add a **Domain Classifier** branch to the `fusion_mlp` to detect the source dataset.
* Train the backbone to produce domain-invariant features.
* **AdaLN (Adaptive Layer Normalization):**
* Switch standard `LayerNorm` to domain-aware normalization.
* Store separate affine parameters ($\gamma, \beta$) for MPII and Gaze360.
### Phase 3: Cross-Domain Distillation
* **Multi-Domain Teacher:** Use L2CS-Net (Pre-trained on Gaze360) to provide soft labels for Gaze360 images.
* **Hybrid Loss:** Combine `Adaptive Wing Loss` (Hard labels) with `KL-Divergence` (Teacher soft labels) across both datasets.
### Phase 4: Self-Supervised Consistency
* **Mirror Invariance:** Enforce $Yaw(Image) = -Yaw(Flipped\_Image)$.
* **Equivariance Loss:** If an image is rotated by $\theta$, the predicted gaze must shift by $\theta$ in the 3D spherical space.
## 4. Architectural Modifications (`src/models/student.py`)
- [ ] Add `GRL` (Gradient Reversal Layer) class.
- [ ] Implement `DomainClassifier` head (2-layer MLP).
- [ ] Update `forward()` to handle `domain_label` and return `domain_logits`.
- [ ] Refactor `geo_mlp` to include `AdaLN` logic.
## 5. Success Criteria
* **Accuracy:** MAE on Gaze360 Frontal subset < 7.0°.
* **Generalization:** Maintain MAE on MPIIGaze < 5.0°.
* **Efficiency:** Latency < 1.5ms on single-core CPU.
---
*Created: 2026-06-04*
*Status: INITIALIZED*