Format and compatibility
New archives use format v7. SQZ and sqz-extract read every earlier version, and the format is documented well enough that an independent reader was written from the specification alone.
Versions at a glance#
| Version | What it added |
|---|---|
| v7 (current) | Folders (including empty ones), symlinks, hard links, permission bits and sub-second times; required-feature bits so older readers refuse cleanly; snapshots stored as changes and chained indexes (spec revision 4) |
| v6 | The x86 branch filter for programs |
| v5 | Full 256-bit BLAKE3 identities for every piece, instead of shorter prefixes |
| v4 | JPEG recompression and the model-family path |
| v1 to v3 | Earlier layouts, still readable |
Who can read what#
- Current SQZ and
sqz-extractread v1 to v7. - New archives need a current build. Archives written now are v7, and those that use chained indexes need a build that knows that feature. SQZ 0.9.1 releases and older cannot open them.
- A reader that meets something newer refuses cleanly. It reports the archive's version, the range it can read, and every feature bit it does not know.
sqz-extractexits with code3in that case.
sqz-extract --version and sqz-extract --info archive.sqz show the versions and features on each side.
Appending to older archives#
New snapshots can only be added to v7 archives. Appending to an older one is refused before anything is changed. There is no in-place migration. To continue an older archive:
sqz x old.sqz C:\old-contents
sqz c C:\old-contents new.sqz
Or keep an older SQZ build around for appending to archives it wrote.
The specification#
The on-disk format is described in a full specification: the header and footer, the index and its required-feature bits, blocks and codecs, piece recipes, whole-file transforms, snapshots and appends, integrity checks, limits, legacy versions, how readers must behave, and the versioning and migration policy. Where the code is the only definition (the context-mixing model, and the preflate and Lepton containers), the specification names the exact source and pinned version.
An independent reader written in Python from the specification alone decodes the store, zstd and LZMA2 codecs, every piece recipe and the x86 filter, and handles chained indexes. It is checked against every frozen test archive and every single-byte corruption of them.
Layout in brief#
offset 0 header 8 bytes: "SQZ2", version, reserved
segment 1 blocks ... | index (zstd) | footer (44 bytes)
segment 2 blocks ... | index (zstd) | footer (after an append)
...
end footer of the newest complete segment
Each append adds a segment. Only the last complete footer counts, so an append that did not finish is simply ignored by readers.
Limits#
| Item | Limit |
|---|---|
| Block, uncompressed | 1 GiB |
| File size for whole-file transforms | 64 MiB |
| JPEG dimensions for recompression | 4096 × 4096 |
| Counts of blocks, pieces, snapshots and files | 232 − 1 each |
| Chained indexes followed from one footer | 64 (writers start a full index after 16) |
Backup repositories#
Backup repositories have their own format (repository version 1), separate from .sqz. It is still a preview and may change. See Security model for how it is built.