|
| static ParquetDiffResult | compare (const std::string &file_1, const std::string &file_2, const ParquetDiffSettings &settings) |
| | Compare two Parquet files.
|
| |
| static ParquetDiffResult | validate (const std::string &file, const std::string &view, const ParquetDiffSettings &settings) |
| | Check one Parquet file against a built-in QPX schema.
|
| |
| static bool | dumpToTsv (const std::string &file, const std::string &out_file, const ParquetDiffSettings &settings) |
| | Write a Parquet table out as TSV, one row per line, sorted by primary key.
|
| |
| static std::vector< std::string > | qpxPrimaryKey (const std::string &view) |
| | The primary key of a QPX view.
|
| |
| static std::vector< std::string > | qpxIdentityColumns (const std::string &view) |
| | The opaque identity and cross-reference columns of a QPX view.
|
| |
| static std::string | viewFromFileType (const std::string &file_type) |
| | Map a QPX file_type metadata value to a view name.
|
| |
Primary-key-aware, order-insensitive, tolerant comparison of Parquet tables.
Two Parquet files holding the same logical table may legitimately differ in row order and in the low-order bits of floating-point columns, so neither a byte comparison nor a text diff of a serialised form answers "are these the same result?". This class matches rows by a primary key, compares the matched rows cell-by-cell with a numeric tolerance, and reports schema drift separately from value drift.
It is the table-shaped counterpart of FuzzyDiff, and is exposed as ParquetDiff.
- Note
- All Arrow/Parquet API use is confined to the implementation, so binaries linking libOpenMS do not need to import Arrow symbols (see ArrowIOHelpers.h).
| static bool dumpToTsv |
( |
const std::string & |
file, |
|
|
const std::string & |
out_file, |
|
|
const ParquetDiffSettings & |
settings |
|
) |
| |
|
static |
Write a Parquet table out as TSV, one row per line, sorted by primary key.
A Parquet file cannot be reviewed in a diff or patched by hand, so a committed binary reference can only ever be regenerated wholesale - which is exactly the operation that hides unrelated drift. Dumping to text puts Parquet output back on the same footing as every other reference in the suite: readable in a pull request, comparable with FuzzyDiff, and editable line by line.
Rows are emitted in primary-key order rather than file order, so the dump does not depend on the order the producer happened to write - the same property that makes the comparison order-insensitive. List- and struct-valued cells are rendered in full; nulls print as null, which is distinct from an empty string or a zero.
- Parameters
-
| [in] | file | the Parquet file to read |
| [in] | out_file | destination TSV path |
| [in] | settings | only ParquetDiffSettings::primary_key is used; when empty the key is derived from the file's QPX file_type metadata, and failing that rows are emitted in file order |
- Returns
- true when the file was read and written successfully