JeonghyeokDo's picture
Upload folder using huggingface_hub
bf48bd4 verified
|
Raw
History Blame Contribute Delete
9.51 kB
# E3Diff — SAR2Opt
**Efficient End-to-End Diffusion Model for One-Step SAR-to-Optical
Translation**, Qin, Zou, Li and Zhang, *IEEE Geoscience and Remote Sensing
Letters*, 2024
([doi:10.1109/LGRS.2024.3506566](https://doi.org/10.1109/LGRS.2024.3506566); the
year is the one in the authors' own citation block — a table that labels this row
with a later issue year is referring to the same paper). Upstream code:
[`DeepSARRS/E3Diff`](https://github.com/DeepSARRS/E3Diff) at commit
`38601093ab8f8e4b478144621f20890b100a3b74`.
This is the authors' two-stage method, **retrained by us end to end** on
SAR2Opt, and it is the checkpoint behind the E3Diff row of the ReFlowSET
comparison table.
**These are our weights, not the authors' released checkpoint.** The authors'
own weights exist for a different dataset and are not what this row measures.
## What is in this folder
| file | bytes | what it is |
|---|---|---|
| `gen.pth` | 768,426,944 | the `sr3` UNet after stage 2, at absolute iteration 310,000 |
Architecturally this is **the same network as the DDPM (SR3-class) row** in this
repository — E3Diff's two stages share one UNet. What differs is the training:
* **stage 1** (250,000 updates) is the eps-prediction conditional DDPM. Its
checkpoint is the `ddpm/` folder in this repository.
* **stage 2** (60,000 further updates, absolute iteration 310,000) fine-tunes
that same network into a **one-step generator**: the sampler is run *with
gradients* from pure noise for `ddim_steps = 1`, and the loss is taken directly
on pixels.
UNet: `inner_channel` 64, `channel_multiplier` [1, 2, 4, 8, 16], `res_blocks` 1,
`norm_groups` 16, no attention, 3 in / 3 out, 3-channel condition.
The condition is the same three-channel `[PPB, Canny, SAR]` stack described on
the DDPM card: FAST_PPB speckle filtering (Deledalle 2009, P = 3, W = 10,
h = 0.5), `cv2.Canny(ppb, 50, 150, L2gradient=True)`, and the raw SAR image.
**This row cannot be run from a SAR PNG alone.**
## Training budget we used
| | |
|---|---|
| **generator updates released** | **250,000 (stage 1, inherited) + 60,000 (stage 2) = 310,000 absolute** |
| batch size / resolution | 4 @ 512 px |
| optimizer | Adam, lr 5e-5 |
| stage-2 sampler during training | DDIM, **1 step**, run with gradients from pure noise |
| stage-2 losses | L1 on pixels, + LPIPS (weight 5), + focal-frequency (weight 10), + a vision-aided CLIP GAN (λ_gan 0.5) |
| stage-1 losses | plain eps-prediction MSE (all auxiliary weights zero) |
| sampler (test) | **DDIM, 1 step** |
| EMA | decay 0.9999 (stage 1) |
| augmentation | flips, 90° rotations, and a random brightness jitter on the SAR/PPB pair (upstream's `transform_augment`) |
The stage-2 budget is 24 % of the stage-1 budget, which follows the authors' own
ratio (they resume at 640,000 and run to 800,000). Note that `n_iter` in this
code base is an **absolute** step count continuing stage 1, not a stage-2
budget — a config that reads 310,000 buys 60,000 new iterations.
## Measured on the SAR2Opt test set (n = 627, 512 px)
| FID↓ | DISTS↓ | LPIPS↓ | SSIM↑ | PSNR↑ |
|---|---|---|---|---|
| 104.7 | 0.232 | 0.529 | 0.249 | 16.09 |
Evaluated on the official split's 627 test tiles, centre-cropped to 512 px.
The tiles are 600 px natively; this benchmark crops and never resizes, in every
method's training and in the evaluation. No subsampling: every metric on this
page is measured over all 627 pairs.
These are our own re-evaluation numbers, measured by us on the images this
checkpoint produced. **No number here is copied from any paper.** PSNR and SSIM
are per-image torchmetrics with `data_range=1`; LPIPS is **LPIPS-VGG on inputs
mapped to [-1, 1]** (the `normalize=False` convention); DISTS is the standard
implementation on the same pairs; FID is `pytorch-fid` against the size-matched
ground-truth test set. The LPIPS convention matters: the alternative [0, 1]
convention gives a systematically different number and the two must never be
mixed, or compared against a paper that used the other one.
## Load it and translate one SAR image
Same entry point and same mechanics as the DDPM row, with `"stage": 2` and
`ddim_steps: 1`.
```python
# 1. Build the condition channels for each SAR image, once:
# ppb = FAST_PPB(sar, P=3, W=10, h=0.5)
# canny = cv2.Canny(ppb.astype('uint8'), 50, 150, L2gradient=True)
# written to <root>/val/SAR-PPB/<name> and <root>/val/SAR-canny/<name>,
# alongside <root>/val/SAR/<name> and <root>/val/EO/<name>.
# 2. Run phase 'val' with a config whose path.resume_state is the checkpoint
# PREFIX -- no '_gen.pth' suffix -- and place gen.pth as <prefix>_gen.pth:
# "stage": 2, "ddim_steps": 1,
# "model": {"beta_schedule": {"val": {"n_timestep": 1, "ddim": 1, ...}}},
# "datasets": {"val": {"data_len": -1, "r_resolution": 512}}
# python main.py -c <val.json> -p val -enable_wandb "" --seed 1
```
Sampling is genuinely one network evaluation per image, which is the point of
the method.
As on the DDPM card, the RGB loader is the one thing you must sort out: upstream
hard-codes a grayscale EO target and a two-channel condition, and with three
channels its `ddim_sample()` slices the condition and concatenates mismatched
shapes. We patched `SAR2EODataset.__getitem__` in memory to return
`HR = EO[0:3]`, `LR = SAR[0:3]`, `SR = cat(PPB[0:1], Canny[0:1], SAR[0:1])` and
then ran the repository's `main.py` verbatim.
## Read before using this checkpoint
* **This is not an oracle.** Worth stating, because a one-step generator invites
the question: the sampler's `condition_x` is the `[PPB, Canny, SAR]` stack, and
the ground truth never enters it. The cell also passes our leak audit on both
datasets.
* **Stage 2 needs the *real* vision-aided CLIP discriminator.** Stage 1 has
`lambda_gan = 0` and can import a stub; stage 2 cannot. If a constant-output
stub shadows the real `vision_aided_loss` package, stage 2 trains against a
constant GAN loss and **silently is not the authors' method**. Assert on
`vision_aided_loss.__file__` before starting a run.
* **SoftPool.** The code imports the SoftPool CUDA extension unconditionally. We
used a pure-PyTorch drop-in with identical maths
(`softpool(x) = avgpool(x·eˣ) / avgpool(eˣ)`) and identical autograd, which
needs no build step.
* **The checkpoint loads with `strict=False`.** A mismatched configuration
**loads nothing and raises no error**. Verify the `Loading pretrained model for
G [...]` log line and look at the first output.
* **`-p val` writes next to the checkpoint and then renames** the sample
directory to `<prefix>_S<ssim>_P<psnr>_l2<l2>_Lp<lpips>`. That rename raises if
the target already exists — after the whole inference has been paid for.
* **`-enable_wandb ""` is mandatory**; the flag defaults to the truthy string
`'false'`.
* One change we made to the repository: `core/logger.py` no longer
unconditionally overwrites `CUDA_VISIBLE_DEVICES` from the config's `gpu_ids`,
which on a shared machine remapped every job onto another user's device. It is
a scheduling fix and does not affect the model.
## Licence — stated factually ⚠ no upstream licence exists
**The upstream code base publishes no licence.**
[`DeepSARRS/E3Diff`](https://github.com/DeepSARRS/E3Diff) has no LICENSE,
LICENCE, COPYING or NOTICE file at the repository root or at any depth; **1 of
its 58 tracked files** is licence-shaped, and it is `SoftPool/LICENSE.txt`, the
MIT licence of a **vendored third-party dependency** (`Copyright (c) 2020
Alexandros Stergiou`, reproduced here as `licenses/LICENSE-SoftPool-MIT.txt`) —
not a grant covering E3Diff. Its README has no licence section, and the GitHub
API reports no declared licence, with the `/license` endpoint returning 404.
Checked 2026-08-28.
Under default copyright that means **all rights are reserved by the authors and
no express permission to redistribute derived work has been granted** to us or
to you. We publish this checkpoint anyway, so that the benchmark is reproducible
end to end, and we state the position plainly rather than implying a permission
that does not exist. If you intend to redistribute this checkpoint or build on
it, assess that for yourself, and consider asking the authors directly.
**Lineage.** E3Diff's README credits SR3
(`Janspiry/Image-Super-Resolution-via-Iterative-Refinement`, Apache-2.0),
`GaParmar/img2img-turbo` (MIT) and `alexandrosstergiou/SoftPool` (MIT). Those
licences cover the borrowed parts only, not the authors' own two-stage
contribution. We ship only the SoftPool text, because it is the only one of the
three that is vendored in the tree we trained from.
Please cite: Qin, Zou, Li and Zhang, *Efficient End-to-End Diffusion Model for
One-Step SAR-to-Optical Translation*, IEEE Geoscience and Remote Sensing
Letters, [doi:10.1109/LGRS.2024.3506566](https://doi.org/10.1109/LGRS.2024.3506566).
The full record of what we checked, per method, is in
`licenses/NO-UPSTREAM-LICENSE.md` at the root of this repository.
---
Part of the **ReFlowSET** release. This folder holds one comparison-method
checkpoint that we retrained ourselves on SAR2Opt; it is not ReFlowSET
itself. Every comparison method in this repository was retrained by us on the
same splits at the same resolution and scored through one evaluation pipeline,
so the rows are directly comparable to each other — and, for the same reason,
**not** directly comparable with the numbers in the methods' own papers.