E2A2 — Encoding Efficient Asymmetric Autoencoders
Trained codecs from github.com/danjacobellis/E2A2. Entropy coding is GGDLPC. Every rate below is actual entropy-coded bits, round-trip verified. Inputs are in [−1, 1]; the encoder quantizer is SC6 softsign companding (coder alphabet [−31, 31]); operating points are channel-prefix truncations of one encode, each decoded with its own merged decoder.
Each codec directory holds config.json, weights.safetensors (merged.{op}.encoders.{s}.*, merged.{op}.decoder.*), entropy.json (the GGDLPC blob), and rd.json (the validation RD curve). Load with the e2a2 package:
from e2a2 import hub
model, coder, config = hub.load_from_hub("audio", n_ch=21) # one operating point
| directory | codec | data | ps / group sizes | stream mode, macroregion | operating points (cumulative channels) |
|---|---|---|---|---|---|
image |
E2A2I, RGB images | LSDIR train / Kodak eval | 32/16/8/4/2 · 3/6/6/6/6 | two_stream, 64 px | 3, 9, 15, 21, 27 |
audio |
E2A2A, 44.1 kHz stereo music | musdb18 (stem remix) | 256/256/128/64/32/16/8 · 4/5/6/6/12/9/9 | all_detail, 16384 samples | 4, 9, 15, 21, 33, 42, 51 |
clean_speech |
16 kHz mono speech, clean→clean | LibriSpeech (clean) | 256/128/64/32/16 · 4/6/6/12/8 | all_detail, 256 samples (16 ms) | 4, 10, 16, 28, 36 |
denoising_speech |
16 kHz mono speech, noisy→clean | same speech + DNS noise + RIRs | 256/128/64/32 · 4/8/8/12 | all_detail, 256 samples (16 ms) | 4, 12, 20, 32 |
hyperspectral |
E2A2H, 224-band AVIRIS radiance as 224 image channels (x = int16 / 32768) |
danjacobellis/aviris_1k / aviris_1k_val |
16/8/4/2/1 · 8/24/24/12/8 | two_stream, 64 px | 8, 32, 56, 68, 76 |
hyperspectral_3d |
the same AVIRIS cubes as a 1-channel 3-d volume (dim = 3, band axis first) |
same | 8/4/2 · 6/6/6 (cubes) | two_stream, 32 voxels | 6, 12, 18 |
spatial_audio |
E2A2S, 7-channel 48 kHz Aria microphone array, raw amplitude (no normalization) | danjacobellis/aria_ea_audio_preprocessed |
256 · 64 (single group; decoder k9/dim 1024/depth 12) | all_detail, 16384 samples | 64 |
kspace |
E2A2K, raw multi-coil MRI k-space, coded as per-coil complex coil images (dim = 3, 2 real channels; the pre-transform — de-companding, line crop to the sampled k-space lines, per-slice inverse FFT, per-item-rms softsign compander with knee 32 — is E2A2_kspace/kspace_img.py in the E2A2 repo and is recorded in config.json under input_domain) |
danjacobellis/fastmri_1e11 (private; the 16-bit codes of github.com/danjacobellis/fastmri) |
8/4/2 · 6/6/6 (cubes) | two_stream, 8 voxels | 6, 12, 18 |
Validation RD (PSNR on the [0, 1]-mapped signal; rates per input sample — bits/pixel for images, bits/frame for audio; rd.json holds each codec's whole-split curve):
hyperspectral (94 tiles, bits per pixel over 224 bands) |
hyperspectral_3d (94 tiles, bits per voxel) |
spatial_audio (1215 clips, bits/frame over 7 ch) |
kspace (161 volumes, 7,707 RSS slices; bits per original complex voxel; display PSNR / SSIM of the p99.5-normalized RSS slices, compressors.fastmri; rd.json also holds LPIPS, DISTS and the linear k-space SNR) |
|---|---|---|---|
| 8: 0.0403, 16.67 dB | 6: 0.0161, 18.49 dB | 64: 0.1063 (5.1 kbps), 31.64 dB | 6: 0.0123, 13.45 dB / 0.196 |
| 32: 0.2631, 17.09 dB | 12: 0.0719, 19.93 dB | 12: 0.2264, 15.64 dB / 0.323 | |
| 56: 0.7716, 17.68 dB | 18: 0.9189, 35.02 dB | 18: 2.3624, 21.50 dB / 0.637 | |
| 68: 4.0749, 18.12 dB | |||
| 76: 5.7383, 18.19 dB |
image (Kodak) |
audio (musdb val, bits/frame) |
clean_speech (bits/sample, kbps; LibriTTS test-clean protocol) |
denoising_speech (vs clean reference) |
|---|---|---|---|
| 3: 0.0107 bpp, 21.54 dB | 4: 0.0320, 28.06 dB | 4: 0.0537 (0.86 kbps), 28.91 dB | 4: 0.0461 (0.74 kbps), 28.81 dB |
| 9: 0.0625 bpp, 25.08 dB | 9: 0.0566, 29.81 dB | 10: 0.1604 (2.57 kbps), 33.72 dB | 12: 0.1653 (2.65 kbps), 32.19 dB |
| 15: 0.2489 bpp, 28.70 dB | 15: 0.1061, 31.24 dB | 16: 0.2942 (4.71 kbps), 36.44 dB | 20: 0.2936 (4.70 kbps), 33.17 dB |
| 21: 0.7541 bpp, 32.94 dB | 21: 0.3196, 33.88 dB | 28: 1.0718 (17.15 kbps), 42.67 dB | 32: 0.6231 (9.97 kbps), 33.76 dB |
| 27: 3.0311 bpp, 40.76 dB | 33: 1.0858, 39.37 dB | 36: 1.6434 (26.30 kbps), 46.43 dB | |
| 42: 1.7583, 41.87 dB | |||
| 51: 3.6039, 45.76 dB |
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support