# Domain Adaptation Plan: LIPE V2 (MPIIGaze → Gaze360) ## 1. Executive Summary The goal of this plan is to bridge the "Domain Gap" between controlled laboratory data (MPIIGaze) and in-the-wild data (Gaze360). Our baseline zero-shot evaluation showed an MAE of **24.3°** on Gaze360. We aim to reduce this to **< 7.0°** while maintaining a CPU-optimized inference speed of > 400 FPS. ## 2. Technical Strategies & Predicted Impact | Phase | Strategy | Technical Mechanism | Predicted MAE | | :--- | :--- | :--- | :---: | | 0 | **Baseline** | Zero-shot from MPIIGaze | 24.3° | | 1 | **Domain Randomization** | Resolution Dropout + Noise Augmentation | 18.5° | | 2 | **Structural Adaptation** | AdaLN + Feature-Level Alignment (DANN) | 12.5° | | 3 | **Expert Guidance** | Multi-Domain Knowledge Distillation | 8.5° | | 4 | **Geometry Constraints** | Self-Supervised Consistency Loss | **< 7.0°** | ## 3. Detailed Implementation Roadmap ### Phase 1: Data Hardening (Augmentation & Preprocessing) * **Resolution Dropout:** Randomly downsample 16x16 patches to 8x8 during training to simulate camera distance. Use `cv2.INTER_CUBIC` to prevent upsampling artifacts. * **Edge-Preserving Denoising:** Integrate **Bilateral Filter** (`cv2.bilateralFilter`) to smooth homogeneous regions and reduce grain without blurring the critical iris boundaries. * **Tuned CLAHE Preprocessing:** Apply CLAHE with a conservative **clipLimit (1.0 - 1.2)** to enhance contrast in dark regions without amplifying noise into artifacts. * **Interpolation Strategy:** Force `INTER_CUBIC` for all resizing to preserve the "Gaze Gradient" in low-resolution conditions. * **Geometric Jitter:** Add Gaussian noise to MediaPipe landmarks (±2px) to simulate tracking instability. * **Extreme Pose Augmentation:** Use existing rotation logic to synthetically create head poses beyond ±20°. ### Phase 2: Feature Alignment (DANN & AdaLN) * **DANN (Domain-Adversarial Neural Network):** * Implement a **Gradient Reversal Layer (GRL)**. * Add a **Domain Classifier** branch to the `fusion_mlp` to detect the source dataset. * Train the backbone to produce domain-invariant features. * **AdaLN (Adaptive Layer Normalization):** * Switch standard `LayerNorm` to domain-aware normalization. * Store separate affine parameters ($\gamma, \beta$) for MPII and Gaze360. ### Phase 3: Cross-Domain Distillation * **Multi-Domain Teacher:** Use L2CS-Net (Pre-trained on Gaze360) to provide soft labels for Gaze360 images. * **Hybrid Loss:** Combine `Adaptive Wing Loss` (Hard labels) with `KL-Divergence` (Teacher soft labels) across both datasets. ### Phase 4: Self-Supervised Consistency * **Mirror Invariance:** Enforce $Yaw(Image) = -Yaw(Flipped\_Image)$. * **Equivariance Loss:** If an image is rotated by $\theta$, the predicted gaze must shift by $\theta$ in the 3D spherical space. ## 4. Architectural Modifications (`src/models/student.py`) - [ ] Add `GRL` (Gradient Reversal Layer) class. - [ ] Implement `DomainClassifier` head (2-layer MLP). - [ ] Update `forward()` to handle `domain_label` and return `domain_logits`. - [ ] Refactor `geo_mlp` to include `AdaLN` logic. ## 5. Success Criteria * **Accuracy:** MAE on Gaze360 Frontal subset < 7.0°. * **Generalization:** Maintain MAE on MPIIGaze < 5.0°. * **Efficiency:** Latency < 1.5ms on single-core CPU. --- *Created: 2026-06-04* *Status: INITIALIZED*