IRENE on-device — ExecuTorch per-step cells (fp16, XNNPACK)

Mobile inference graphs for IRENE, the ConvGRU ensemble precipitation-nowcasting model (64M params, © FBK DSIP, BSD-2-Clause). Exported from the released model.ckpt as per-timestep cells: the caller drives both recurrences in a loop, carrying the five per-block ConvGRU hidden states — so the forecast rollout length and the input-history length are loop counts, not export parameters. One encoder/decoder pair per domain size.

Produced by export_executorch.py --cells --fp16 in the irene-android repo (wrappers and the full I/O contract in its export/common.py; parity gate in export/parity_step.py).

Graph contract (num_blocks=5 ⇒ divisor 32)

graph inputs outputs
encoder_step frame (1,1,H,W) normalized reflectivity + states e1 (1,1,H,W) … e5 (1,256,H/16,W/16) (zeros at t=0) e1'..e5' + h1..h5 (downsampled block outputs, deepest first — the last step's are the decoder's initial states)
decoder_step noise_t (1,1024,H/32,W/32) ~N(0,1) + states h1..h5 frame_t (1,1,H,W) normalized, unclamped + h1'..h5'

All I/O is fp16 (no in-graph casts). Pre/post-processing lives in the caller: pad to /32 → NaN→0 → Marshall–Palmer dBZ = 10·log10(200·R^1.6+1e-16) clip [0,60] → /30−1; inverse + crop on the way out. Per ensemble member: re-seed the decoder states from the encoder skips and draw fresh noise per step. Trained config: 6 input frames, 12 forecast steps (5-min cadence); longer rollouts run mechanically but are beyond the trained horizon.

Two integration warnings (cost us a day each):

  • On Android, Module.forward output tensors alias the runtime's memory arena — copy every carried state before the next call.
  • Pin the ExecuTorch runtime to the exporting version: 1.4.0 (delegate blobs do not cross versions).

Files

file raw zst
encoder_step_{256x256,512x512,1024x1024,1408x1216}_fp16.pte 15.1 MB each 7.9 MB each
decoder_step_{256x256,512x512,1024x1024,1408x1216}_fp16.pte 241.6 MB each 122.1 MB each

1408x1216 is the full Italian DPC mosaic (1400×1200, padded to /32); the others are crop sizes in km (1 km cells).

Note on size: the weights inside every file are the same 64M fp16 parameters; they occupy 2× their fp16 bytes because ExecuTorch ≤1.4 serializes convolution filters as fp32 for XNNPACK (op_conv2d.py force_fp32=True — the f16 conv API takes fp32 filters and packs at load). That's also why zstd compresses them almost exactly 2:1 (the upcast halves have zero low-mantissa bits), and why the per-size patches below are tiny: the weight bytes are identical across sizes, only the graph skeleton differs.

Getting the files

Each .pte.zst is standalone:

zstd -d decoder_step_512x512_fp16.pte.zst
sha256sum -c SHA256SUMS.pte

Minimal-bandwidth set (~130 MB for all four sizes): download the two 256 files + patches-from-256/, then reconstruct:

zstd -d --long=28 --patch-from=decoder_step_256x256_fp16.pte \
  decoder_step_1024x1024_fp16.pte.zst-patch -o decoder_step_1024x1024_fp16.pte

(decoder patches ~218 KB, encoder patches ~2 KB; all reconstructions are byte-identical, verify with SHA256SUMS.pte.)

patches-from-512/ is the same idea against the 512 base — kept for irene-android ≤0.5.9, whose in-app updater patches from 512.

Validation

  • Eager cells ≡ the monolithic export ≡ RadarLightningModel.predict() bit-exact (fixed noise, rain-rate space).
  • ExecuTorch fp32 cells vs eager: ~1e-4 mm/h mean abs.
  • These fp16 cells produce output bit-identical to the monolithic fp16 export; fp16-vs-fp32 error measured on-device at 0.079 mm/h mean against a 7.2 mm/h ensemble spread.
  • Measured on a Snapdragon 8 Gen 2 (4 threads, mmap): 512² ≈ 206 ms/step, full grid ≈ 1.3 s/step — a 10-member 1-hour ensemble in ~25 s, a full-grid 2-hour member in ~40 s.

License

BSD-2-Clause, © Fondazione Bruno Kessler (DSIP) — same as the IRENE model these are derived from.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DSIP-FBK/irene-mobile

Base model

it4lia/irene
Quantized
(1)
this model