SSQZ/atlasDocs
Archives

Modes and tuning

Every mode deduplicates, streams with bounded memory and verifies every file. They differ in how hard they look for a smaller result, and so in how long packing and unpacking take.

The four modes#

ModeWhat it does per blockUnpacks mediaGood for
--fast
alias --instant
zstd level 3NoQuick packing and transfer, data that is already compressed
--smartA planner predicts the size of store, zstd-3, zstd-9 and LZMA2 from a sample and weighs it against their speed; close calls are compressed in fullYesTrying the new planner; the default in the Windows menu
--balancedCompresses with LZMA2 and zstd and keeps the smallerYesEveryday archives
--max
default, alias --maximum
Races SQZ's context-mixing engine against LZMA2 on samples, then compares the winner with zstdYesLong-term archives where size matters more than time

"Unpacks media" means the reversible recompression of PNG, ZIP, Office, PDF, gzip and JPEG files described below.

What to expect#

A real Windows app build, 158.4 MB, packed on a Windows PC in September 2026 (one run):

ModeArchiveTime to pack
--fast88.1 MB0.4 s
--balanced35.1 MB19.4 s
--max31.3 MB57.4 s
ZIP, for comparison91 MB
7z, for comparison79 MBslower than Balanced

On the Python 3.11 standard library (51.4 MB, 4 threads), the archive was 31.6% of the input with Fast in 0.3 s, 25.9% with Smart in 9.3 s, 23.2% with Balanced in 10.1 s and 19.2% with Max in 46 s. Your data will differ; sqz estimate tells you for your own files, and sqz bench compares all four modes on them.

Choosing#

The engine inside Max#

Max can use SQZ's own context-mixing (CM) engine, written in C++. It predicts each bit from many models at once and does best on text, source code, tables and program files. Against xz -9e (the LZMA2 setting 7-Zip Ultra uses) it was 15% to 30% smaller on source code, CSV data and program files, and tied on an ELF symbol table. It runs at about 0.5 MB/s per thread and uses about 260 MB per thread, which is why Max is slow and why it only uses CM where a sample shows it wins.

Threads#

sqz c --balanced -j 4 C:\data data.sqz

-j N (or --threads N) sets how many blocks are compressed at once. Fast uses every core by default. Balanced, Smart and Max use up to 8, because each thread can need hundreds of MiB. Fewer threads means less memory and a slower job. Extraction also decodes blocks in parallel and takes -j too.

On Windows 11 with hybrid CPUs, SQZ opts out of power throttling so it is not confined to the efficiency cores.

Block size#

sqz c --max --block 128 C:\data data.sqz

--block MiB sets the largest solid block, from 4 to 1024 MiB (default 32). Bigger blocks can compress large files better, because the codec sees more data at once, but they use more memory and give fewer blocks to spread over threads.

Reversible recompression#

Most files people store are already compressed inside: PNG images, ZIP files, Word, Excel and PowerPoint documents (which are ZIP files), PDFs, gzip files and JPEG photos. Compressing them again gains almost nothing. In Balanced, Smart and Max, SQZ opens them up instead:

A transformed file is accepted only if SQZ has rebuilt the original bytes exactly, including any trailing data, and the result is smaller (for JPEG, by at least 32 bytes). Otherwise the file is stored normally. Whole-file transforms apply to files up to 64 MiB, and JPEGs up to 4096×4096 pixels.

Ten PNG screenshots packed to 52.0% of their size with Balanced and 46.8% with Max, against 92% for LZMA2 alone.

--no-transforms turns this off, mainly for benchmark comparisons. sqz probe <file-or-folder> shows, for each transform, how many bytes it saves and what it costs to encode and decode.

AI model checkpoints#

sqz c --fast --model-family -j 2 C:\models family.sqz

--model-family is an experimental path for folders of related .safetensors checkpoints (a base model and its fine-tunes, for example). It deduplicates identical tensor pieces, rearranges BF16, FP16 and FP32 values into layouts that compress better, and stores exact XOR differences against a matching tensor from an earlier checkpoint in the same archive. Every bit pattern, NaN payloads and signed zeros included, comes back exactly. Decoding needs no GPU, no download and no floating-point arithmetic. Put the base and related checkpoints in the same input folder. Unsupported data falls back to ordinary compression, and pickle files are never deserialized.