OpenMS
Loading...
Searching...
No Matches
ParquetDiff

Compares two Parquet tables by primary key, tolerating numeric differences.

The table counterpart of FuzzyDiff. Two Parquet files holding the same logical result may legitimately differ in row order and in the low-order bits of floating-point columns, so neither a byte comparison nor a text diff answers "is this the same result?". ParquetDiff matches rows by a primary key, compares matched rows cell by cell with a numeric tolerance, and reports schema drift separately from value drift.

Only one of 'ratio' or 'absdiff' has to be satisfied. Use "absdiff" to deal with cases like "zero vs. epsilon".

Rows are matched on 'pk'. For a QPX file (psm, feature or pg) the key is derived from the file's own file_type metadata when 'pk' is not given. List-valued key columns are canonicalised by sorting their elements, so a set-valued key such as grouped_runs matches regardless of the order the producer emitted.

With 'schema' the tool takes a single input and checks it against the built-in QPX schema for that view instead of comparing two files; this reports missing or extra columns, wrong Arrow types, wrong nullability, and duplicate primary keys.

'min_rows' requires every input to hold at least that many rows, and is checked in both modes. It closes a gap neither of them covers: two empty tables compare equal, and an empty table satisfies any schema, so a producer that wrote a correctly shaped nothing passes both. Combining schema with min_rows therefore asserts something about a single file without needing a reference at all - useful where a reference would only add churn. '0' is accepted and always passes; use it to record deliberately that a table is expected to be empty.

The command line parameters of this tool are:

ParquetDiff -- Compares two Parquet tables by primary key, tolerating numeric differences.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ParquetDiff.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
 + Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
   trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.

Usage:
  ParquetDiff <options>

Options (mandatory options marked with '*'):
                     
  -in1 <file>*       First input file (valid formats: 'parquet', 'pqt')
  -in2 <file>        Second input file (omit when 'schema' is given) (valid formats: 'parquet', 'pqt')
                     
  -pk <string list>  Primary-key columns used to match rows. If empty, derived from the QPX 'file_type' metad
                     ata of the first file.
  -schema <view>     Instead of comparing two files, check 'in1' against the built-in QPX schema of this view
                      (valid: 'psm', 'feature', 'pg')
  -out_tsv <file>    Instead of comparing, write 'in1' out as TSV sorted by primary key. A Parquet reference 
                     can only be regenerated wholesale, which hides unrelated drift; a text dump can be revie
                     wed in a diff and patched line by line, and compared with FuzzyDiff like every other 
                     reference. (valid formats: 'tsv')
                     
  -ratio <double>    Acceptable relative error. Only one of 'ratio' or 'absdiff' has to be satisfied.  Use 
                     "absdiff" to deal with cases like "zero vs. epsilon". (default: '1.0') (min: '1.0')
  -absdiff <double>  Acceptable absolute difference. Only one of 'ratio' or 'absdiff' has to be satisfied.  
                     (default: '0.0') (min: '0.0')
                     
  -min_rows <int>    Require every input to hold at least this many rows ('-1' does not check). A schema chec
                     k alone is satisfied by an empty table, so this is what distinguishes 'the tool wrote a 
                     correct table' from 'the tool wrote a correctly shaped nothing'. '0' is accepted and 
                     always passes: use it to record that a table is expected to be empty. (default: '-1') 
                     (min: '-1')
                     
  -verbose <int>     Set verbose level:
                     0 = very quiet mode (absolutely no output)
                     1 = quiet mode (no output unless differences detected)
                     2 = default (include summary at end)
                      (default: '2') (min: '0' max: '2')
                     
Common TOPP options:
  -ini <file>        Use the given TOPP INI file
  -threads <n>       Sets the number of threads allowed to be used by the TOPP tool (0 = all available cores)
                      (default: '1')
  -write_ini <file>  Writes the default configuration file
  --help             Shows options
  --helphelp         Shows all options (including advanced)

INI file documentation of this tool:

Legend:
required parameter
advanced parameter

This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.

Each parameter entry contains the following information:

  • Name The identifier used in configuration files and on the command line.
  • Default value The value used if the parameter is not explicitly specified.
  • Description A short explanation describing the purpose and behavior of the parameter.
  • Tags Additional metadata associated with the parameter.
  • Restrictions Allowed value ranges for numeric parameters or valid options for string parameters.

Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.

Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.

+ParquetDiffCompares two Parquet tables by primary key, tolerating numeric differences.
version3.6.0-pre-nightly-2026-09-29 Version of the tool that generated this parameters file.
++1Instance '1' section for 'ParquetDiff'
in1 first input fileinput file*.parquet, *.pqt
in2 second input file (omit when 'schema' is given)input file*.parquet, *.pqt
pk[] primary-key columns used to match rows. If empty, derived from the QPX 'file_type' metadata of the first file.
schema instead of comparing two files, check 'in1' against the built-in QPX schema of this viewpsm, feature, pg
out_tsv instead of comparing, write 'in1' out as TSV sorted by primary key. A Parquet reference can only be regenerated wholesale, which hides unrelated drift; a text dump can be reviewed in a diff and patched line by line, and compared with FuzzyDiff like every other reference.output file*.tsv
ratio1.0 acceptable relative error. Only one of 'ratio' or 'absdiff' has to be satisfied. Use "absdiff" to deal with cases like "zero vs. epsilon".1.0:∞
absdiff0.0 acceptable absolute difference. Only one of 'ratio' or 'absdiff' has to be satisfied. 0.0:∞
min_rows-1 require every input to hold at least this many rows ('-1' does not check). A schema check alone is satisfied by an empty table, so this is what distinguishes 'the tool wrote a correct table' from 'the tool wrote a correctly shaped nothing'. '0' is accepted and always passes: use it to record that a table is expected to be empty.-1:∞
ignore[] columns excluded from value comparison. They are still schema-checked.
schema_onlyfalse compare schemas only; do not compare valuestrue, false
with_idsfalse also compare the QPX identity columns (feature_id/psm_id/pg_id and the cross-references between them). Off by default: those values are derived, and the spec states that identity is meaningful within a file only, so comparing them across two files is a join it tells you not to make. 'schema' mode checks instead that they are present, non-null and unique.true, false
unordered_listsfalse compare list-valued cells as multisets rather than sequencestrue, false
max_reported25 stop listing differences of one kind after this many (0 = unlimited)0:∞
verbose2 set verbose level:
0 = very quiet mode (absolutely no output)
1 = quiet mode (no output unless differences detected)
2 = default (include summary at end)
0:2
log Name of log file (created only when specified)
debug0 Sets the debug level
threads1 Sets the number of threads allowed to be used by the TOPP tool (0 = all available cores)
no_progressfalse Disables progress logging to command linetrue, false
forcefalse Overrides tool-specific checkstrue, false
testfalse Enables the test mode (needed for internal use only)true, false