phi-JEPA / README.md
kiki-orb's picture
Update README.md
b3ed31c verified
|
Raw
History Blame Contribute Delete
2.76 kB
metadata
library_name: pytorch
pipeline_tag: feature-extraction
tags:
  - synthetic-aperture-radar
  - sar
  - remote-sensing
  - self-supervised-learning
  - masked-image-modeling
  - jepa
  - vision-transformer

phi-JEPA: Where and What to Reconstruct in Physics-Informed Radar Pre-training?

This is the official model repository for phi-JEPA, a physics-informed masked autoencoder for self-supervised synthetic aperture radar (SAR) image pre-training.

Release status. This repository currently provides the pretrained checkpoint used in our main experiments. The pre-training code and downstream evaluation code are not included in the current release.

Model Description

Existing masked image modeling methods for SAR imagery commonly rely on random masking and low-level reconstruction targets. phi-JEPA introduces two SAR-oriented designs:

  1. Scattering-Center-Aware Masked Modeling (SCM). A physical scattering prior reconstructed from attributed scattering centers guides masking toward target-related regions while maintaining a high overall masking ratio.
  2. SAR-Specific Semantic Feature Prediction (SFP). A frozen SARCLIP ViT-B/16 image encoder provides patch-level semantic targets, shifting the pre-training objective from low-level appearance reconstruction to semantic representation learning.

The released ViT-B/16 checkpoint is pretrained on ATRNet-STAR and is intended for downstream SAR representation transfer.

Released Checkpoint

Model Backbone Pre-training dataset Epochs File SHA-256
phi-JEPA ViT-B/16 ATRNet-STAR 100 phi_jepa_vitb16_e100.pth f94f048366528815c15a51a060ea27d923db65ae2c6434931121cfd7af72f095

Download

Using the Hugging Face CLI:

hf download kiki-orb/phi-JEPA \
  main_experiments/phi_jepa_vitb16_e100.pth \
  --local-dir ./checkpoints

Using Python:

from huggingface_hub import hf_hub_download

checkpoint_path = hf_hub_download(
    repo_id="kiki-orb/phi-JEPA",
    filename="main_experiments/phi_jepa_vitb16_e100.pth",
)
print(checkpoint_path)

Pre-training Configuration

Configuration Value
Backbone ViT-Base
Patch size 16 x 16
Pre-training dataset ATRNet-STAR (SOC, 40 vehicle categories)
Optimizer AdamW
Base learning rate 1e-3
Weight decay 0.05
Batch size 128
Pre-training epochs 100
Warmup 20 epochs
Learning-rate schedule Cosine annealing
Overall masking ratio 80%
Target-region masking ratio 50%
Feature alignment objective Cosine loss
Semantic teacher Frozen SARCLIP ViT-B/16 image encoder