Buckets:

1.79 GB
339 files
Updated 6 days ago
Name
Size
decompressor
README.md506 Bytes
xet
archive.xz24.6 MB
xet
decompressor.zip848 Bytes
xet
enwik8.out100 MB
xet
preprocessed.bin86.9 MB
xet
results.json371 Bytes
xet
run_experiment.py3.59 kB
xet
README.md

dict-auto-xz (AutoZip)

Approach:

  • Detect bytes not present in enwik8.
  • Substitute high-gain XML/wiki/text patterns with those single-byte codes.
  • Compress the transformed stream with tuned xz/lzma2.
  • Decompress by xz decode + inverse substitution table (subs.json).

Backend xz options:

  • -9e --lzma2=dict=512MiB,nice=273,mf=bt4,mode=normal,lc=4,lp=0,pb=0

Roundtrip:

  • Verified byte-identical output against shared_resources/enwik8.

Score:

  • See results.json (archive + decompressor.zip).
Total size
1.79 GB
Files
339
Last updated
Aug 12
Pre-warmed CDN
US EU US EU

Contributors

  • +7