IRENE on-device — ExecuTorch per-step cells (fp16, XNNPACK)
Mobile inference graphs for IRENE, the
ConvGRU ensemble precipitation-nowcasting model (64M params, © FBK DSIP,
BSD-2-Clause). Exported from the released model.ckpt as per-timestep
cells: the caller drives both recurrences in a loop, carrying the five
per-block ConvGRU hidden states — so the forecast rollout length and the
input-history length are loop counts, not export parameters. One
encoder/decoder pair per domain size.
Produced by export_executorch.py --cells --fp16 in the irene-android repo
(wrappers and the full I/O contract in its export/common.py; parity gate in
export/parity_step.py).
Graph contract (num_blocks=5 ⇒ divisor 32)
| graph | inputs | outputs |
|---|---|---|
encoder_step |
frame (1,1,H,W) normalized reflectivity + states e1 (1,1,H,W) … e5 (1,256,H/16,W/16) (zeros at t=0) |
e1'..e5' + h1..h5 (downsampled block outputs, deepest first — the last step's are the decoder's initial states) |
decoder_step |
noise_t (1,1024,H/32,W/32) ~N(0,1) + states h1..h5 |
frame_t (1,1,H,W) normalized, unclamped + h1'..h5' |
All I/O is fp16 (no in-graph casts). Pre/post-processing lives in the
caller: pad to /32 → NaN→0 → Marshall–Palmer dBZ = 10·log10(200·R^1.6+1e-16)
clip [0,60] → /30−1; inverse + crop on the way out. Per ensemble member:
re-seed the decoder states from the encoder skips and draw fresh noise per
step. Trained config: 6 input frames, 12 forecast steps (5-min cadence);
longer rollouts run mechanically but are beyond the trained horizon.
Two integration warnings (cost us a day each):
- On Android,
Module.forwardoutput tensors alias the runtime's memory arena — copy every carried state before the next call. - Pin the ExecuTorch runtime to the exporting version: 1.4.0 (delegate blobs do not cross versions).
Files
| file | raw | zst |
|---|---|---|
encoder_step_{256x256,512x512,1024x1024,1408x1216}_fp16.pte |
15.1 MB each | 7.9 MB each |
decoder_step_{256x256,512x512,1024x1024,1408x1216}_fp16.pte |
241.6 MB each | 122.1 MB each |
1408x1216 is the full Italian DPC mosaic (1400×1200, padded to /32); the
others are crop sizes in km (1 km cells).
Note on size: the weights inside every file are the same 64M fp16
parameters; they occupy 2× their fp16 bytes because ExecuTorch ≤1.4
serializes convolution filters as fp32 for XNNPACK (op_conv2d.py
force_fp32=True — the f16 conv API takes fp32 filters and packs at load).
That's also why zstd compresses them almost exactly 2:1 (the upcast halves
have zero low-mantissa bits), and why the per-size patches below are
tiny: the weight bytes are identical across sizes, only the graph skeleton
differs.
Getting the files
Each .pte.zst is standalone:
zstd -d decoder_step_512x512_fp16.pte.zst
sha256sum -c SHA256SUMS.pte
Minimal-bandwidth set (~130 MB for all four sizes): download the two 256
files + patches-from-256/, then reconstruct:
zstd -d --long=28 --patch-from=decoder_step_256x256_fp16.pte \
decoder_step_1024x1024_fp16.pte.zst-patch -o decoder_step_1024x1024_fp16.pte
(decoder patches ~218 KB, encoder patches ~2 KB; all reconstructions are
byte-identical, verify with SHA256SUMS.pte.)
patches-from-512/ is the same idea against the 512 base — kept for
irene-android ≤0.5.9, whose in-app updater patches from 512.
Validation
- Eager cells ≡ the monolithic export ≡
RadarLightningModel.predict()bit-exact (fixed noise, rain-rate space). - ExecuTorch fp32 cells vs eager: ~1e-4 mm/h mean abs.
- These fp16 cells produce output bit-identical to the monolithic fp16 export; fp16-vs-fp32 error measured on-device at 0.079 mm/h mean against a 7.2 mm/h ensemble spread.
- Measured on a Snapdragon 8 Gen 2 (4 threads, mmap): 512² ≈ 206 ms/step, full grid ≈ 1.3 s/step — a 10-member 1-hour ensemble in ~25 s, a full-grid 2-hour member in ~40 s.
License
BSD-2-Clause, © Fondazione Bruno Kessler (DSIP) — same as the IRENE model these are derived from.
- Downloads last month
- 11
Model tree for DSIP-FBK/irene-mobile
Base model
it4lia/irene