Export ConsensusMap feature data to Apache Arrow format following QPX feature schema.
More...
#include <OpenMS/FORMAT/ConsensusMapArrowExport.h>
Export ConsensusMap feature data to Apache Arrow format following QPX feature schema.
This class provides static methods to export ConsensusMap data to Apache Arrow Tables and Parquet files. The schema follows the QPX (Quantitative Proteomics Exchange) feature format.
- Experimental classes:
- This API is experimental and may change in future versions.
◆ exportToArrow()
Export ConsensusMap to Apache Arrow Table.
Exports consensus features following the QPX feature schema. Each ConsensusFeature becomes one row with identification, quantification, and protein group information.
- Parameters
-
| [in] | cmap | The ConsensusMap to export |
| [out] | out_links | Optional feature↔PSM linkage, collected while the rows are built. Pass it to the psm exporter to fill psm.feature_id. Producing it here rather than in a second pass is what keeps the two directions reciprocal – qpx validates that a PSM pointing at a feature is listed back by that feature's psm_ids. The export refuses a psm_id claimed by two different feature rows instead of silently choosing one owner. |
- Returns
- Shared pointer to Arrow Table, or nullptr on error
- Exceptions
-
◆ exportToParquet()
Export ConsensusMap to Parquet file.
- Parameters
-
| [in] | cmap | The ConsensusMap to export |
| [in] | filename | Output file path |
| [in] | config | Parquet writing options |
| [out] | out_links | Optional feature↔PSM linkage, see exportToArrow() |
- Returns
- true on success, false on error
- Exceptions
-
◆ exportToParquetStreaming()
Stream a ConsensusMap to a Parquet file in row batches (bounded peak memory)
Functionally equivalent to exportToParquet() but builds and flushes the feature table one batch_size -sized range at a time through a persistent parquet::arrow::FileWriter, instead of materializing the whole ~N-row Arrow table in memory before a single write. For isobaric data (one consensus feature per PSM) N can be in the millions, where the one-shot path's transient peak drives the process into swap / OOM; here peak memory stays bounded by one batch.
Each batch is optionally partitioned and built in parallel with OpenMP and written in index order (the Parquet writer stays serial), so the written rows and their order are identical to exportToParquet() and deterministic for any thread count; only the Parquet row-group layout may differ.
- Parameters
-
| [in] | cmap | The ConsensusMap to export |
| [in] | filename | Output file path |
| [in] | batch_size | Consensus features materialized per batch (0 is treated as the default) |
| [in] | config | Parquet writing options |
| [in] | n_threads | OpenMP threads for the per-batch build: 1 = serial (default), 0 = all available cores (honors OMP_NUM_THREADS), N = fixed |
| [out] | out_links | Optional feature↔PSM linkage, see exportToArrow(). Each worker fills its own map and they are merged in index order, so the collected linkage does not depend on n_threads either. |
- Returns
- true on success, false on error
◆ requireResolvableIdRuns()
| static void requireResolvableIdRuns |
( |
const ConsensusMap & |
cmap | ) |
|
|
static |
Refuse a map whose features cannot be attributed to an origin MS run.
A feature's identification names its run through id_merge_index into the identification run's spectra_data. Without a usable index every PSM of a merged run resolves to the run's FIRST file, so run_file_name would be wrong rather than missing.
Only the identification the exported row actually uses is validated – validating every attached one refuses a feature whose winning hit resolves perfectly because a sibling, hitless or simply not selected, lacks an index that is never read.
- Note
- Same preflight constraint as requireUnambiguousIdentities(): call before any OpenMP region and before the output file is opened.
- Parameters
-
| [in] | cmap | The map about to be exported |
- Exceptions
-
◆ requireUnambiguousIdentities()
| static void requireUnambiguousIdentities |
( |
const ConsensusMap & |
cmap | ) |
|
|
static |
Verify that every feature can be attributed to a single peptide.
The QPX feature and pg views have singular sequence / peptidoform / charge and no ConsensusFeature identifier, so a feature whose identifications disagree on the top peptide cannot be represented: exporting one would publish a chosen interpretation while silently discarding the alternatives, and the pg view would count evidence contributed by one peptide under another's identity. The ambiguity-preserving representation is the OpenMS-native consensusparquet (ConsensusMapArrowIO), which keeps every hit and the consensus feature association, and which IDConflictResolver accepts as input.
Several identifications agreeing on the same modified sequence (FEATURE_ID_MULTIPLE_SAME) are unambiguous and accepted – that is the deliberate output of IDConflictResolverAlgorithm::resolve() with keep_matching.
- Note
- Must be called as a preflight, before any OpenMP region and before the output file is opened. An exception escaping the parallel batch build would be flattened into a boolean by the worker's catch-all, and one escaping the region at all is std::terminate.
- Parameters
-
| [in] | cmap | The map about to be exported |
- Exceptions
-