SSQZ/atlasDocs
Reference

How it works

You run one command. Behind it, every file takes the same trip through six steps, and every step can be undone exactly.

1. Sort#

SQZ scans the input and lines files up by type, so similar files sit next to each other. Text goes near text, programs near programs, which helps the compressor find patterns across files.

2. Unpack#

Files that are already compressed inside (PNG, ZIP and Office files, PDF, gzip, JPEG) are opened up, so the real content can be compressed better. This happens only if SQZ can rebuild the original bytes exactly, and it proves that before accepting the result. x86 and x86-64 programs go through a branch filter that turns relative jump targets into absolute ones, which repeat more. Fast mode skips this step. See Reversible recompression.

3. Chunk#

Data is cut into pieces of 16 to 256 KiB (about 64 KiB on average) at points chosen by the content itself, not at fixed offsets. A small edit only changes the pieces around it, and the same content produces the same pieces wherever it appears.

4. Skip repeats#

Each piece gets a full 256-bit BLAKE3 fingerprint. A piece that is already stored, from this file, another file, or an earlier snapshot, is kept once and pointed to. An image embedded in two documents, or many versions of one file, is stored once.

5. Group#

New pieces are routed into groups (text, program code, binary, raw and images) because like compresses best with like. Groups are packed into solid blocks of 4 to 32 MiB by default (--block changes the ceiling).

6. Squeeze and seal#

Each block is compressed with the codec your mode picks, and blocks are compressed in parallel. Each block records only the ID of the codec used, so a reader never needs the planner that chose it. Finally, the index (which files, which pieces, which blocks, which snapshots) is compressed, fingerprinted and written with a footer at the end of the file.

The codecs#

CodecCharacterUsed by
storeNo compression, for data that will not shrinkSmart; any mode when compression would grow a block
zstdVery fast both ways, moderate sizeFast (level 3), Smart (3 or 9), and as a comparison in Balanced and Max
LZMA2Slow to pack, fairly fast to unpack, smallSmart, Balanced, Max
SQCMSQZ's own context-mixing engine: slowest both ways, smallest on text and codeMax, where a sample shows it wins

Extraction#

Extraction plans every piece it needs up front, decodes each block once, in parallel, and keeps only the pieces still needed later, within a fixed memory budget. When many files share deduplicated data this matters: a 162 MB Windows app set with 36 MB of shared data unpacks in 144 s instead of 475 s on 4 cores. Every piece is checked against its fingerprint and every file against its digest before it is published. See Integrity and safe restores.

Appending#

An append (sqz a) repeats the pipeline for the new snapshot. Pieces already in the archive are pointed to, new blocks are written after the existing data, and a new index describing only what changed is written with a new footer. See Snapshots.

And backups?#

sqz backup uses the same ideas with different trade-offs: larger pieces (about 1 MiB, with boundaries keyed per repository), zstd compression, encryption of every object, and a repository of immutable files instead of one growing file. See Backup basics.