Data Compression¶
Compression can considerably reduce disk space usage and the time needed for data transfer.
However, compressing and decompressing data consumes CPU and memory resources, can slow down read and write operations, and bring potential loss of information depending on the chosen compression strategy.
When to compress
Ideally, compress data that is not accessed frequently. Compression is highly recommended for data that is stored in an archive (see Archiving Data).
Lossless vs. lossy compression¶
| Lossless | Lossy | |
|---|---|---|
| Principle | Restores the file to its original state, without the loss of a single bit | Permanently removes bits that are redundant, unimportant, or imperceptible |
| Typical use | Measured and simulated data, archives | Graphics, audio, video, images; depending on the use case also measured or simulated data |
| Reversible | Yes | No |
| File size | Smaller | Usually smallest |
Working with compressed files¶
Compressed files such as *.gz, *.zip, or *.bz2 normally need to be uncompressed before they can be used again. There are alternatives, however:
- Linux commands such as
zless,zcat,zdiff, andzgrephandle compressed files on the fly. - many packages from scripting languages such as R, Python, Matlab, or IDL can read compressed files directly.
Standard lossless methods¶
Recommendation
We recommend lossless compressed netCDF4 files for most use cases. The netCDF4 library, and all tools compiled with netCDF4 support, integrate compression seamlessly.
Lossless compression in netCDF4 is based on the zlib library. The following parameters can be tuned:
- Compression level: ranges from 1 (least aggressive) to 9 (most aggressive). Level 1 requires moderate CPU and memory resources and is sufficient for most purposes. Higher levels usually produce smaller files, but at a higher processing cost when packing and unpacking and often for a marginal additional gain.
- Chunk sizes: compression operates on data chunks. These should match the data blocks that are typically accessed at the same time.
- Shuffling: often further reduces the data size.
- Unlimited dimensions: remove unneeded unlimited dimensions, as they may reduce the compression efficiency.
Tools¶
Part of every netCDF4 installation. Allows setting the compression level, the shuffling option, and chunk sizes.
# Compression level 1 with shuffling
nccopy -d 1 -s input.nc output.nc
# Additionally set chunk sizes per dimension
nccopy -d 1 -s -c time/1,lat/180,lon/360 input.nc output.nc
The ncks command is part of the NCO toolkit and offers extensive chunking options.
ncks -4 -L 1 input.nc output.nc
IAC systems
On IAC systems, the script nczip is installed. It is a wrapper around ncks and compresses or decompresses a single netCDF file or all netCDF files in a folder.
Climate Data Operators can compress netCDF files, but offer fewer chunking options than NCO.
cdo -f nc4 -z zip_1 copy input.nc output.nc
nccompress is another alternative to nczip for batch compression of netCDF files.
Compression settings can also be specified directly in the source code when creating netCDF variables, e.g., via the Fortran, R, or Python interfaces.
import xarray as xr
ds = xr.open_dataset("input.nc")
ds.to_netcdf("output.nc", encoding={"temp": {"zlib": True, "complevel": 1, "shuffle": True}})
Lossy algorithms¶
Use with care
Lossy compression permanently removes information that cannot be restored. Always verify that the error introduced is acceptable for your use case before archiving or sharing data.
Lossy compression can achieve ratios of 10–50× for climate and weather data, far beyond what lossless methods offer. Available tools include:
-
NCO (ncks)
In addition to lossless compression (see Standard lossless methods),
nckssupports several lossy algorithms. -
Python
netcdf4-python and xarray support lossy compression by defining a least significant digit. All information below this threshold is removed before saving.
-
C2SM dc_toolkit
The data-compression toolkit systematically searches for the best compression pipeline for your data and verifies the result against a user-defined error threshold. See C2SM data-compression toolkit below.
C2SM data-compression toolkit (dc_toolkit)¶
The dc_toolkit automates the search for the best compression pipeline for netCDF files and writes the result into a zarr zip. It sweeps all compressor × filter × serialiser combinations on a representative sample of each field and filters results according to pre-defined error thresholds.
Key concepts¶
Error metrics:
- L1 (mean absolute error): average deviation across all cells; the most common budget for climate data.
- L2 (root-mean-square error): penalises large local deviations.
- L∞ (maximum absolute error): the worst-case deviation anywhere in the field.
Compression pipeline — compression is applied in up to three stages:
| Stage | Role | Examples |
|---|---|---|
| Serialiser | Converts floating-point values into bytes | zfp, EBCC, FixedScaleOffset |
| Filter | Pre-processes data to improve compressibility | Delta, BitRound, AsType |
| Compressor | Applies a byte-level codec | Zstd, Blosc, LZ4 |
Note: Community recommendations — the ESiWACE3 project publishes community recommendations specifying the maximum permissible error for ERA5 variables. These can be used as the gate of a sweep as a replacement for manual error thresholds.
Workflow¶
-
evaluate_combos— sweeps the codec space on a sample of each field and writes the best pipeline into a json file:dc_toolkit evaluate_combos input.nc \ --where-to-write ./path_to_folder \ --field-to-compress field \ --l1-threshold 0.005 \ # relative L1 error budget (0.5 %) --eval-data-size-limit 5GB -
compress— writes all fields into a shared.zarrstore and re-reads every field to verify the stored data meets the original thresholds:dc_toolkit compress input.nc ./out
Compression libraries¶
- Numcodecs : the standard Zarr codec library, providing Zstd, Blosc, LZ4, FixedScaleOffset, Delta, etc.
- EBCC (optional): the Error Bounded Climate Compressor. At loose error bounds (0.1–1 % of the field's range) it achieves 2–4× the ratio of
zfp; at tight bounds the advantage disappears. However, it is slow to encode (~1–2 MB/s per core).
Installation¶
git clone https://github.com/C2SM/data-compression.git dc_toolkit
cd dc_toolkit
python -m venv venv && source venv/bin/activate
bash install_dc_toolkit.sh
HPC usage¶
On santis, load the prgenv-gnu/26.3:v1 uenv first. On balfrin, use netcdf-tools/2024:v1.
#SBATCH --nodes=8 --ntasks-per-node=32 --cpus-per-task=1
export OMP_NUM_THREADS=1 MKL_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 \
BLOSC_NTHREADS=1 NUMBA_NUM_THREADS=1 \
VECLIB_MAXIMUM_THREADS=1 OMP_THREAD_LIMIT=1
srun dc_toolkit evaluate_combos input.nc \
--where-to-write ./out \
--field-to-compress t \
--l1-threshold 0.005 \
--eval-data-size-limit 5GB
For further details, including Docker usage and a step-by-step walkthrough, see the dc_toolkit documentation .