Compresses and uncompresses byte buffers using Zstandard (zstd), optionally combined with the byte-shuffling and dictionary transforms defined for mzML.
More...
|
| static void | compressData (const void *raw_data, size_t in_length, std::string &compressed_data, int level=DEFAULT_LEVEL) |
| | Compress the in_length bytes pointed to by raw_data into compressed_data using zstd.
|
| |
| static void | uncompressData (const void *compressed_data, size_t nr_bytes, std::string &out, size_t expected_size=0) |
| | Uncompress the zstd-compressed compressed_data.
|
| |
| static void | byteShuffle (const void *data, size_t nr_bytes, size_t element_size, std::string &out) |
| | Byte-shuffle an array of element_size byte elements.
|
| |
| static void | byteUnshuffle (const void *data, size_t nr_bytes, size_t element_size, std::string &out) |
| | Reverse the byte shuffling done by byteShuffle().
|
| |
| static void | dictionaryEncode (const void *data, size_t nr_bytes, size_t element_size, std::string &out) |
| | Dictionary-encode an array of little-endian element_size byte elements (values and indices byte-shuffled).
|
| |
| static void | dictionaryDecode (const void *data, size_t nr_bytes, size_t element_size, std::string &out, size_t array_length=0) |
| | Decode a buffer created by dictionaryEncode() back into an array of little-endian element_size byte elements.
|
| |
| static void | encode (const void *data, size_t nr_bytes, ByteTransform transform, size_t element_size, std::string &out, int level=DEFAULT_LEVEL) |
| | Apply transform to an array of little-endian element_size byte elements and compress the result with zstd.
|
| |
| static void | decode (const void *data, size_t nr_bytes, ByteTransform transform, size_t element_size, std::string &out, size_t array_length=0) |
| | Uncompress zstd data and reverse transform, yielding an array of little-endian element_size byte elements.
|
| |
Compresses and uncompresses byte buffers using Zstandard (zstd), optionally combined with the byte-shuffling and dictionary transforms defined for mzML.
Static utility class implementing the zstd based binary data array compression methods recommended by the mzML specification (PSI-MS CV terms MS:1003780 to MS:1003785):
- zstd compression (MS:1003780): the little-endian bytes of the array are compressed with zstd.
- byte-shuffled zstd compression (MS:1003781): the bytes of the array elements are transposed ("shuffled") so that the i-th byte of every element is stored contiguously, followed by zstd compression. This usually compresses sorted data (e.g. m/z arrays) much better.
- dictionary-encoded zstd compression (MS:1003782): the array is replaced by a sorted dictionary of its distinct values and an array of indices into that dictionary. Values and indices are byte-shuffled separately, then everything is compressed with zstd. This is well suited for arrays with many repeated values (e.g. ion mobility, charge states).
- The MS-Numpress variants (MS:1003783 to MS:1003785) apply plain zstd compression to the output of the respective MS-Numpress encoder.
The layout of a dictionary-encoded buffer is: an unsigned 64 bit little-endian integer holding the byte offset of the index array, an unsigned 64 bit little-endian integer holding the number of distinct values n, the byte-shuffled sorted distinct values, and finally the byte-shuffled indices. The indices use the smallest unsigned integer type able to address all values (8 bit if n < 2^8, 16 bit if n < 2^16, 32 bit if n < 2^32, otherwise 64 bit). Since implementations disagree on the index width at these boundaries (e.g. 8 bit indices for exactly 256 values), the decoder derives the width from the size of the index region if the number of array elements is known (see dictionaryDecode()).
All array transforms operate on raw byte buffers holding a contiguous array of fixed-size elements in little-endian byte order (the byte order mandated by mzML). The data is treated as raw bytes and may contain embedded zeros.
- See also
- https://github.com/mobiusklein/mzd.cpp for the reference implementation.
| static void dictionaryDecode |
( |
const void * |
data, |
|
|
size_t |
nr_bytes, |
|
|
size_t |
element_size, |
|
|
std::string & |
out, |
|
|
size_t |
array_length = 0 |
|
) |
| |
|
static |
Decode a buffer created by dictionaryEncode() back into an array of little-endian element_size byte elements.
An empty input yields an empty output.
If array_length is given, the index width is taken from the size of the index region divided by array_length, provided that this is 1, 2, 4 or 8 bytes and adjacent to the width defined by the specification. This accepts buffers written with the diverging boundary conventions of other implementations (e.g. 8 bit indices for 256 values or 16 bit indices for 255 values). Otherwise, the width defined by the specification is used.
- Parameters
-
| [in] | data | Pointer to the dictionary-encoded bytes. |
| [in] | nr_bytes | Length of data in bytes. |
| [in] | element_size | Size of a single array element in bytes (1, 2, 4 or 8). |
| [out] | out | Receives the decoded array bytes; any previous contents are replaced. |
| [in] | array_length | Number of elements of the encoded array, e.g. the mzML arrayLength (0 if unknown). |
- Exceptions
-
| static void uncompressData |
( |
const void * |
compressed_data, |
|
|
size_t |
nr_bytes, |
|
|
std::string & |
out, |
|
|
size_t |
expected_size = 0 |
|
) |
| |
|
static |
Uncompress the zstd-compressed compressed_data.
Multiple concatenated frames and frames without a recorded content size are supported. An empty input yields an empty output.
The output buffer is initially sized from the content size recorded in the frame header, capped at expected_size (or, if that is 0, at a small multiple of nr_bytes), and grows as needed. The expected size therefore only affects the initial allocation, not the result.
- Parameters
-
| [in] | compressed_data | Pointer to the zstd-compressed bytes. |
| [in] | nr_bytes | Length of compressed_data in bytes. |
| [out] | out | Receives the decompressed bytes; any previous contents are replaced. |
| [in] | expected_size | Expected size of the decompressed data in bytes (0 if unknown). |
- Exceptions
-