ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks
Paper • 1809.00219 • Published
A 4x image upscaler model based on the ESRGAN architecture, trained from scratch on DIV2K + Flickr2K.
| Model | Purpose | FP16 | FP32 |
|---|---|---|---|
| Forensic | When fidelity matters more than looks | forensic_fp16.safetensors | forensic_fp32.safetensors |
| Perceptual | Photos, art, anything for viewing | perceptual_fp16.safetensors | perceptual_fp32.safetensors |
Standard ESRGAN architecture (RRDB backbone, 4x upscale). Two-stage training:
Trained on DIV2K + Flickr2K with bicubic 4x downsampling for the LR side. Total finetune length: 685k iterations.
This is a perceptual-quality SR model, not a forensic one. It's good at producing plausible high-resolution output, not at recovering ground truth that isn't there.
The forensic variant exists for cases where you'd rather have honest softness than confident invention.