Estimates
sqz estimate packs a sample of your real data and tells you how long a big job will take, how large the archive will be and how much memory it needs, before you commit to it.
Before packing#
sqz estimate [--fast|--smart|--balanced|--max] [-j N] <file-or-folder> [--json]
sqz estimate --max D:\projects
sqz estimate --balanced -j 4 D:\projects --json
SQZ compresses a sample with the mode you choose, on one thread: whole files up to 4 MiB and 1 MiB slices of larger ones, spread evenly over the input. The sample is up to 32 MiB for Fast, 16 for Smart, 12 for Balanced and 8 for Max, so the estimate itself takes from a few seconds to half a minute. If the whole input fits in the sample, the size is measured rather than estimated.
Before extracting#
sqz estimate [-s N] [-j N] [--file PATH] <archive.sqz> [<output-folder>]
sqz estimate big.sqz D:\restore
sqz x --estimate big.sqz D:\restore
For an archive, SQZ finds the blocks the chosen files need, decodes up to 32 MiB of them and times each codec. The output size is exact, from the index. Give an output folder and it also reports the free space there.
Automatic estimates#
sqz c and sqz a print an estimate by themselves when the input is 1 GiB or more. --no-estimate skips it, and --estimate forces it on c, a and x.
How accurate it is#
Measured against real runs on 4 cores and 16 GB of memory, one run each (estimate divided by actual; 1.00 is exact):
| Data | Mode | Size | Time | Memory |
|---|---|---|---|---|
| Rust sources, 310 MB | Balanced | 1.14 | 0.67 | 0.99 |
| Rust sources, 310 MB | Max | 1.22 | 1.00 | 1.04 |
| /usr/share, 278 MB | Balanced | 0.97 | 0.80 | 1.02 |
| Build output, 422 MB | Max | 1.77 | 1.35 | 0.92 |
| Random + text, 305 MB | Max | 1.06 | 0.86 | 1.11 |
| Browser bundle, 968 MB | Balanced | 1.45 | 0.71 | 0.99 |
- Time for Balanced and Max was within 0.67× to 1.35× of the real run: roughly within a third either way.
- Memory was within 11% of the real peak every time.
- Size is an upper bound. A sample cannot see duplicate files or repeats far apart, so data full of them (build outputs, bundles) ends up smaller than estimated.
- Fast and Smart time is reported as "at least about": for Fast the disk decides the speed, and Smart picks slower codecs where they pay off. Smart memory is reported as the Balanced figure, the most it can use.
Other machines, hybrid CPUs, network drives, very slow disks, thread counts other than the default, extraction estimates on large archives, and Windows.
For scripts#
--json prints the numbers as one line of JSON, so a script can decide whether to go ahead. Both forms include kind ("pack" or "extract"), files, bytes, secs, mem_mib and threads. A packing estimate adds mode, the estimated archive size in bytes, and sampled_bytes, sample_secs and read_secs for the sample itself.
sqz estimate --balanced D:\projects --json