Buckets:
1.79 GB
339 files
Updated 6 days ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| decompressor | 2 items | ||
| README.md | 506 Bytes xet | 99ba2cca | |
| archive.xz | 24.6 MB xet | 41860fa9 | |
| decompressor.zip | 848 Bytes xet | 12cf10d3 | |
| enwik8.out | 100 MB xet | b63296fe | |
| preprocessed.bin | 86.9 MB xet | 860fe4c9 | |
| results.json | 371 Bytes xet | 4aa0e7ac | |
| run_experiment.py | 3.59 kB xet | 144ebd85 |
dict-auto-xz (AutoZip)
Approach:
- Detect bytes not present in enwik8.
- Substitute high-gain XML/wiki/text patterns with those single-byte codes.
- Compress the transformed stream with tuned xz/lzma2.
- Decompress by xz decode + inverse substitution table (
subs.json).
Backend xz options:
-9e --lzma2=dict=512MiB,nice=273,mf=bt4,mode=normal,lc=4,lp=0,pb=0
Roundtrip:
- Verified byte-identical output against
shared_resources/enwik8.
Score:
- See
results.json(archive + decompressor.zip).
- Total size
- 1.79 GB
- Files
- 339
- Last updated
- Aug 12
- Pre-warmed CDN
- US EU US EU