SSTVAE

A hybrid-digital SSTV mode for HF radio: images are sent as the continuous-valued latents of a convolutional autoencoder, modulated onto OFDM carrier amplitudes. The decoder is trained with a simulated HF channel in the loop, so quality degrades gradually with SNR instead of falling off a digital cliff. Inspired by FreeDV RADE.

⚠️ Working beta. Decoded successfully over the air, but the on-air format is not frozen β€” expect incompatible changes. Both stations must run the same code and the same checkpoint; there's no handshake, so a mismatch decodes to noise rather than erroring.

Contents

  • v1.pt β€” the published checkpoint (10.3 M params, 640Γ—480)
  • v1-{encoder,decoder}-{fp32,fp16,int8}.onnx β€” the same codec exported to ONNX, six artifacts (see below)
  • v1-onnx-manifest.json β€” sizes, SHA-256s and measured error for each
  • sstvae/, scripts/train.py, scripts/export_onnx.py, pyproject.toml β€” model, training and export code matching this checkpoint

The command-line tools (encode / decode / live receive) and full documentation live in the project repository, github.com/arodland/SSTVAE, not here. The tools fetch v1.pt from this repo automatically, so there's nothing to download by hand.

ONNX artifacts

Published so that anyone building a compatible receiver can run the codec without torch β€” onnxruntime is about 27 MB against torch's 336 MB, and a receiving station needs nothing else from torch.

Nothing in the current tools loads these yet. codec.py is still torch; the artifacts exist for third-party implementers and for a future runtime path. They are exported from v1.pt by scripts/export_onnx.py at publish time, and each file records the source checkpoint's name and SHA-256 in its metadata_props, so an artifact can always be traced back to the checkpoint it came from.

There is one on-air format

The precisions are not variants of it. Every precision decodes every other precision's transmission β€” fp32 ONNX and torch agree to ~2e-06 on unit-RMS latents, roughly 105 dB below the channel noise. Which file you load is a local packaging choice, invisible to the far end.

All three are published rather than left to implementers, because a third party choosing their own quantisation settings is the case that actually risks divergence.

Which to use

Measured against the torch model on 10 COCO validation images, and separately on non-photographic images β€” the two differ more than you might expect, so both are given:

encoder decoder encoder latent error photos non-photographic
fp32 23.6 MB 17.9 MB 1.95e-06 RMS βˆ’0.000 dB βˆ’0.000 dB
fp16 11.9 MB 9.0 MB 4.58e-04 RMS βˆ’0.000 dB βˆ’0.000 dB
int8 6.5 MB 8.6 MB 7.31e-02 RMS βˆ’0.002 dB βˆ’0.112 dB

fp16 is the recommended default. Half the size of fp32 at no measurable cost. Note it is not faster on the CPU provider β€” the convolutions up-cast to fp32 β€” so treat it as a download-size win.

int8 is also close to free, at 2.7Γ— smaller than fp32. That is not automatic: each int8 artifact keeps its single most quantisation-sensitive convolution at fp32, chosen by measurement at export time. A plain quantize_dynamic of the fp32 artifact costs 0.19 dB on photographs and 1.6 dB on everything else, so prefer these published artifacts to rolling your own.

Encoder and decoder are not the same decision. Decoder precision only changes the picture you see and needs no coordination with anyone. Encoder precision changes what goes on the air, and its error is paid by every receiver β€” the int8 encoder sits 14.0 dB under the channel noise, about 0.17 dB of effective SNR, which is acceptable but not free. fp16 costs nothing at all, so use it unless size really binds.

int8 is slower than fp32 on x86 for these convolution shapes. That is a kernel-matching artifact and typically reverses on ARM; benchmark on your target rather than trusting either result.

Performance

Measured end-to-end through the real modem and a simulated channel, on validation images, PSNR in dB:

mode time clean 10 dB 6 dB 0 dB
A 32 s 24.7 24.4 24.0 22.3
B 64 s 25.7 25.4 25.0 23.5
C 95 s 26.0 25.8 25.5 24.1

~1200 Hz occupied bandwidth, ~4.5 dB envelope PAPR, >Β±50 Hz frequency offset tolerance. Below about βˆ’2 dB SNR the limit is acquisition β€” you get no image rather than a poor one.

Output is a reconstruction, not a photograph: fine detail, small text especially, can come back subtly wrong rather than merely blurry. Don't use it where exactness matters.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support