![]() |
OpenMS
|
Compares two Parquet tables by primary key, tolerating numeric differences.
The table counterpart of FuzzyDiff. Two Parquet files holding the same logical result may legitimately differ in row order and in the low-order bits of floating-point columns, so neither a byte comparison nor a text diff answers "is this the same result?". ParquetDiff matches rows by a primary key, compares matched rows cell by cell with a numeric tolerance, and reports schema drift separately from value drift.
Only one of 'ratio' or 'absdiff' has to be satisfied. Use "absdiff" to deal with cases like "zero vs. epsilon".
Rows are matched on 'pk'. For a QPX file (psm, feature or pg) the key is derived from the file's own file_type metadata when 'pk' is not given. List-valued key columns are canonicalised by sorting their elements, so a set-valued key such as grouped_runs matches regardless of the order the producer emitted.
With 'schema' the tool takes a single input and checks it against the built-in QPX schema for that view instead of comparing two files; this reports missing or extra columns, wrong Arrow types, wrong nullability, and duplicate primary keys.
'min_rows' requires every input to hold at least that many rows, and is checked in both modes. It closes a gap neither of them covers: two empty tables compare equal, and an empty table satisfies any schema, so a producer that wrote a correctly shaped nothing passes both. Combining schema with min_rows therefore asserts something about a single file without needing a reference at all - useful where a reference would only add churn. '0' is accepted and always passes; use it to record deliberately that a table is expected to be empty.
The command line parameters of this tool are:
ParquetDiff -- Compares two Parquet tables by primary key, tolerating numeric differences.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ParquetDiff.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
+ Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.
Usage:
ParquetDiff <options>
Options (mandatory options marked with '*'):
-in1 <file>* First input file (valid formats: 'parquet', 'pqt')
-in2 <file> Second input file (omit when 'schema' is given) (valid formats: 'parquet', 'pqt')
-pk <string list> Primary-key columns used to match rows. If empty, derived from the QPX 'file_type' metad
ata of the first file.
-schema <view> Instead of comparing two files, check 'in1' against the built-in QPX schema of this view
(valid: 'psm', 'feature', 'pg')
-out_tsv <file> Instead of comparing, write 'in1' out as TSV sorted by primary key. A Parquet reference
can only be regenerated wholesale, which hides unrelated drift; a text dump can be revie
wed in a diff and patched line by line, and compared with FuzzyDiff like every other
reference. (valid formats: 'tsv')
-ratio <double> Acceptable relative error. Only one of 'ratio' or 'absdiff' has to be satisfied. Use
"absdiff" to deal with cases like "zero vs. epsilon". (default: '1.0') (min: '1.0')
-absdiff <double> Acceptable absolute difference. Only one of 'ratio' or 'absdiff' has to be satisfied.
(default: '0.0') (min: '0.0')
-min_rows <int> Require every input to hold at least this many rows ('-1' does not check). A schema chec
k alone is satisfied by an empty table, so this is what distinguishes 'the tool wrote a
correct table' from 'the tool wrote a correctly shaped nothing'. '0' is accepted and
always passes: use it to record that a table is expected to be empty. (default: '-1')
(min: '-1')
-verbose <int> Set verbose level:
0 = very quiet mode (absolutely no output)
1 = quiet mode (no output unless differences detected)
2 = default (include summary at end)
(default: '2') (min: '0' max: '2')
Common TOPP options:
-ini <file> Use the given TOPP INI file
-threads <n> Sets the number of threads allowed to be used by the TOPP tool (0 = all available cores)
(default: '1')
-write_ini <file> Writes the default configuration file
--help Shows options
--helphelp Shows all options (including advanced)
INI file documentation of this tool:
This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.
Each parameter entry contains the following information:
Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.
Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.