OpenMS
Loading...
Searching...
No Matches
ProteomicsLFQ

Performs label-free quantification of peptides and proteins.

Input:

  • Spectra in mzML format, Bruker .d directories (TimsTOF PASEF) or Thermo .raw files (read as acquired, without vendor peak picking; see Vendor formats)
  • Identifications in idXML or mzIdentML format with posterior error probabilities as score type. To generate those we suggest to run:

    1. PeptideIndexer to annotate target and decoy information.
    2. PSMFeatureExtractor to annotate percolator features.
    3. PercolatorAdapter tool (score_type = 'q-value', -post_processing_tdc)
    4. IDFilter (pep:score = 0.01) to filter PSMs at 1% FDR

    Exactly one identification run per ID file is required, and merged ID runs are not supported. One identification per spectrum is expected as well: ProteomicsLFQ measures one value per (spectrum, peptidoform, charge), so where several identifications of one spectrum agree on all three, only the best-scoring one is kept and the reduction is reported. The others would otherwise count the same measurement more than once, in the PSM-level FDR and in every output. Results from several search engines must therefore be combined - with ConsensusID (-algorithm best -keep_old_scores, which preserves each engine's score) - rather than simply concatenated. Identifications of one spectrum that name different peptidoforms are left alone: a chimeric spectrum yields two distinct measurements.

  • An experimental design file:
    (see ExperimentalDesign for details)
  • A protein database in with appended decoy sequences in FASTA format
    (e.g., generated by the OpenMS DecoyDatabase tool)
    Processing:
    ProteomicsLFQ has different methods to extract features: ID-based (targeted only), or both ID-based and untargeted.
  1. The first method uses targeted feature dectection using RT and m/z information derived from identification data to extract features. Note: only identifications found in a particular MS run are used to extract features in the same run. No transfer of IDs (match between runs) is performed.
  2. The second method adds untargeted feature detection to obtain quantities from unidentified features. Transfer of Ids (match between runs) is performed by transfering feature identifications to coeluting, unidentified features with similar mass and RT in other runs.

Match between runs with PIP-ECHO (-pip_echo true):
By default, the second method links the features across runs by QT clustering on their RT and m/z (Linking:*), without estimating the error rate of the transferred identifications. With -pip_echo true, PIP-ECHO links the features instead and controls the false discovery rate of the transfers. Decoy transfers, which look for a peptide at the retention time of an unrelated peptide, estimate the error rate. A support vector machine scores each transfer (on intensity, mass error, RT agreement, isotope envelope and, if every run has ion mobility data, ion mobility), and the transfers up to PipEcho:fdr (default 0.05) are kept. -pip_echo requires -targeted_only false (the default).

  • The RT window in which a transfer is searched is local by default: for each peptide, it is predicted from nearby peptides identified in both runs and sized from their RT scatter (PipEcho:local_rt:*), and it is widened if there are too few decoys to resolve the FDR. PipEcho:local_rt:enabled false uses one global window instead, which ProteomicsLFQ sizes from the alignment error and the chromatographic peak width.
  • If there are fewer decoy transfers than PipEcho:min_decoys (default 20), or too few to resolve the requested FDR, no transfer is kept, and only the direct identifications remain. PipEcho:fdr 1.0 keeps all transfers.
  • PipEcho:random_seed selects the decoys; with the same seed (default 0), results are reproducible.
  • PipEcho:max_training_points limits how many transfers the support vector machine is trained on in each cross-validation fold (default 50000, 0 = no limit), which bounds the run time on large data sets. All transfers are scored.

Resuming and distributing feature detection (-feat_dir):
Feature detection is the expensive part of the workflow and each MS run is detected independently of every other; alignment, linking, inference and quantification need all runs at once. -feat_dir names a directory of per-run feature checkpoints and applies one rule to every row of the experimental design: reuse its checkpoint if a valid one is there, otherwise detect the run from -in / -ids and write one, otherwise fail.

Everything follows from that rule:

// one machine, resumable -- interrupt it, run it again, it continues
ProteomicsLFQ -design d.tsv -in *.mzML -ids *.idXML -fasta db.fasta -feat_dir ckpt/ -out r.mzTab
// many machines: one detect-only invocation per run, in any order, no coordination
ProteomicsLFQ -design d.tsv -in a.mzML -ids a.idXML -fasta db.fasta -feat_dir ckpt/ -detect_only
// then combine, reading no mzML, no idXML and no run-level FASTA content at all
ProteomicsLFQ -design d.tsv -fasta db.fasta -feat_dir ckpt/ -out r.mzTab
if none is it is fetched and built automatically from source via CMake FetchContent Adds< code > ProteomicsLFQ
Definition common-cmake-parameters.doxygen:76

A checkpoint records the parameters, OpenMS build, experimental-design row and input files it was produced from, and is refused if any of those disagree with the run trying to use it - naming the setting that differs. This is what makes reuse safe rather than merely convenient: nothing else would stop half a study being detected with one setting and half with another. There is no way to combine checkpoints that disagree: -force_recompute detects the affected runs again and rewrites their checkpoints.

-feat_dir requires an explicit -design (a generated one would label every separately detected run as the first) and -fasta (a checkpoint has to carry the peptide-indexing results, which a combining run cannot reconstruct), and does not apply to spectral_counting.

Note on scale: the combining step holds every run's features of a fraction in memory at once, so its ceiling is set by features per run rather than by run count. Measured on a 7-run dataset averaging about 10,000 features per run, the marginal cost is roughly 2 kB per feature plus 10 MB per run, so 200 such runs in one fraction need about 9 GB; a deep-proteome experiment at ~60,000 features per run would need several times that.

FAIMS (Field Asymmetric Ion Mobility Spectrometry):
FAIMS data is automatically detected based on compensation voltage (CV) annotations in the mzML file. The data is split by CV and processed separately for each voltage group during feature detection. Features representing the same analyte detected at different CV values are merged automatically. The merged features are then aligned and linked across runs based on RT and m/z. No special preparation of the input mzML file is required.

Bruker .d (TimsTOF PASEF):
Bruker .d directories containing DDA-PASEF data are supported directly. When .d input is detected, the tool automatically:

  • Skips centroiding (PeakPickerHiRes) to preserve per-peak ion mobility data
  • Skips precursor mass correction (not IM-aware)
  • Forces Biosaur2Algorithm for seed generation (FeatureFinderMultiplex does not support IM_PEAK)
  • Estimates chromatographic FWHM from Biosaur2 feature extents Identification files should be generated with SageAdapter, which annotates ion mobility values in the idXML output. FeatureFinderIdentificationAlgorithm uses these IM annotations for targeted 2D chromatogram extraction (m/z + IM windowing). MS1 frames are IM-centroided during loading using BrukerTimsFile's built-in Sage algorithm, collapsing ~245k raw peaks/frame into ~10k centroided peaks with summed intensity. Biosaur2 defaults are tuned for ProteomicsLFQ (mini=500, minlh=3, pasefminlh=2). On HeLa 50ng 5-min timsTOF gradient: 34k seeds, 80% model fit success, 2,809 peptides quantified (Spearman r=0.62 vs Sage LFQ), 75s runtime, 1.3 GB memory.

Normalization:

  • ProteomicsLFQ does NOT normalize. Every output - mzTab, consensusXML, QPX Parquet and MSstats - reports the same un-normalized abundances, whatever combination of output files is requested. Earlier versions applied median scaling to the consensus features unless -out_msstats was given, which made the meaning of an intensity depend on which other output file happened to be requested, and recorded nothing about the transform in any of them.
  • Choosing a normalization is a separate, explicit step. Three routes, in increasing distance from this tool:
    • ConsensusMapNormalizer on -out_cxml (median, quantile, robust regression or thresholded scaling);
    • the ProteinQuantification:consensus:normalize parameter, which scales peptide abundances so that the median of each (fraction group, label) assay matches the overall median - i.e. it normalizes at the assay level rather than per fraction;
    • the downstream tool itself. MSstats and comparable consumers normalize their input by design.

Output (at least one required; each output is optional individually):

  • mzTab file with analysis results (out)
  • MSstats file with analysis results for statistical downstream analysis in MSstats (out_msstats)
  • consensusXML file for visualization and further processing in OpenMS (out_cxml)
  • QPX Parquet collection (out_qpx)

The command line parameters of this tool are:

ProteomicsLFQ -- A standard proteomics LFQ pipeline.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ProteomicsLFQ.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
 + Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
   trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.

Usage:
  ProteomicsLFQ <options>

Options (mandatory options marked with '*'):
  -in <file list>                                            Input files. Optional only when '-feat_dir' supp
                                                             lies a checkpoint for every run of the design. 
                                                             (valid formats: 'mzML', 'd', 'raw')
  -ids <file list>                                           Identifications filtered at PSM level (e.g., 
                                                             q-value < 0.01).And annotated with PEP as main 
                                                             score.
                                                             We suggest using:
                                                             1. PSMFeatureExtractor to annotate percolator 
                                                             features.
                                                             2. PercolatorAdapter tool (score_type = 'q-value
                                                             ', -post_processing_tdc)
                                                             ...
                                                             than concatenating them. (valid formats: 'idXML'
                                                             , 'mzId', 'idparquet')
  -design <file>                                             Design file (valid formats: 'tsv')
  -fasta <file>                                              Fasta file (valid formats: 'fasta', 'fa', 'faa')

  -out <file>                                                Optional output mzTab file. At least one output 
                                                             must be specified. (valid formats: 'mzTab')
  -out_msstats <file>                                        Optional output MSstats input file. At least 
                                                             one output must be specified. (valid formats: 
                                                             'csv')
  -out_cxml <file>                                           Optional output consensusXML file. At least one 
                                                             output must be specified. (valid formats: 'conse
                                                             nsusXML')
  -out_qpx <directory>                                       Optional output directory for QPX Parquet files 
                                                             (quantms.feature.parquet, quantms.psm.parquet, 
                                                             quantms.pg.parquet). At least one output must 
                                                             be specified.
  -feat_dir <directory>                                      Directory of per-run feature checkpoints. For 
                                                             every run of the experimental design, a valid 
                                                             checkpoint here is reused instead of detecting 
                                                             features again; a run without one is detected 
                                                             from '-in'/'-ids' and its checkpoint written. 
                                                             This makes a run resumable, and lets the per-run
                                                              work be distributed: run with '-detect_only' 
                                                             on each machine, then once over the design with 
                                                             neither '-in' nor '-ids'. Requires '-design' 
                                                             and '-fasta'.
  -detect_only                                               Stop after the per-run feature checkpoints have 
                                                             been written. No alignment, linking, inference 
                                                             or quantification is performed, and no result 
                                                             file is required. Requires '-feat_dir'.
  -proteinFDR <threshold>                                    Protein FDR threshold (0.05=5%). (default: '0.05
                                                             ') (min: '0.0' max: '1.0')
  -picked_proteinFDR <choice>                                Use a picked protein FDR? (default: 'false') 
                                                             (valid: 'true', 'false')
  -psmFDR <threshold>                                        FDR threshold for sub-protein level (e.g. 0.05=5
                                                             %). Use -FDR_type to choose the level. Cutoff 
                                                             is applied at the highest level. If Bayesian 
                                                             inference was chosen, it is equivalent with a 
                                                             peptide FDR (default: '1.0') (min: '0.0' max: 
                                                             '1.0')
  -FDR_type <threshold>                                      Sub-protein FDR level. PSM, PSM+peptide (best 
                                                             PSM q-value). (default: 'PSM') (valid: 'PSM', 
                                                             'PSM+peptide')
  -quantification_method <option>                            Feature_intensity: MS1 signal.
                                                             spectral_counting: PSM counts. (default: 'featur
                                                             e_intensity') (valid: 'feature_intensity', 'spec
                                                             tral_counting')
  -targeted_only <option>                                    True: Only ID based quantification.
                                                             false: include unidentified features so they 
                                                             can be linked to identified ones (=match between
                                                              runs). (default: 'false') (valid: 'true', 'fals
                                                             e')
  -pip_echo <option>                                         Perform match between runs (MBR) via PIP-ECHO 
                                                             (default: 'false') (valid: 'true', 'false')

Parameters for seeding of untargeted features:
  -Seeding:algorithm <choice>                                Algorithm for untargeted seed feature detection.
                                                             
                                                             multiplex: FeatureFinderMultiplexAlgorithm (defa
                                                             ult, current behavior).
                                                             biosaur2: Biosaur2Algorithm (handles IM_PEAK/PAS
                                                             EF data natively). (default: 'multiplex') (valid
                                                             : 'multiplex', 'biosaur2')

Centroiding:
  -Centroiding:signal_to_noise <value>                       Minimal signal-to-noise ratio for a peak to be 
                                                             picked (0.0 disables SNT estimation!) (default: 
                                                             '0.0') (min: '0.0')
  -Centroiding:ms_levels <numbers>                           List of MS levels for which the peak picking is 
                                                             applied. If empty, auto mode is enabled, all 
                                                             peaks which aren't picked yet will get picked. 
                                                             Other scans are copied to the output without 
                                                             changes. (min: '1')

PeptideQuantification:
  -PeptideQuantification:quantify_decoys                     Whether decoy peptides should be quantified (tru
                                                             e) or skipped (false).
  -PeptideQuantification:min_psm_cutoff <text>               Minimum score for the best PSM of a spectrum to 
                                                             be used as seed. Use 'none' for no cutoff. (defa
                                                             ult: 'none')
  -PeptideQuantification:add_mass_offset_peptides <value>    If for every peptide (or seed) also an offset 
                                                             peptide is extracted (true). Can be used to down
                                                             stream to determine MBR false transfer rates. 
                                                             (0.0 = disabled) (default: '0.0') (min: '0.0')

Parameters for ion chromatogram extraction:
  -PeptideQuantification:extract:batch_size <number>         Nr of peptides used in each batch of chromatogra
                                                             m extraction. Smaller values decrease memory 
                                                             usage but increase runtime. (default: '5000') 
                                                             (min: '1')
  -PeptideQuantification:extract:mz_window <value>           M/z window size for chromatogram extraction (uni
                                                             t: ppm if 1 or greater, else Da/Th) (default: 
                                                             '10.0') (min: '0.0')
  -PeptideQuantification:extract:IM_window <value>           Ion mobility (IM) window for chromatogram extrac
                                                             tion in the IM dimension. Set to 0.0 to disable 
                                                             IM filtering (even if data contains IM informati
                                                             on). The window is applied as +/- IM_window/2 
                                                             around the median IM value of identified peptide
                                                             s. This parameter is automatically ignored if 
                                                             the input data does not contain IM information 
                                                             (determined via IMTypes::determineIMFormat). 
                                                             ...
                                                             for quality control. (default: '0.06') (min: 
                                                             '0.0')

Parameters for detecting features in extracted ion chromatograms:
  -PeptideQuantification:detect:mapping_tolerance <value>    RT tolerance (plus/minus) for mapping peptide 
                                                             IDs to features. Absolute value in seconds if 1 
                                                             or greater, else relative to the RT span of the 
                                                             feature. (default: '0.0') (min: '0.0')

Parameters for fitting exp. mod. Gaussians to mass traces.:
  -PeptideQuantification:EMGScoring:max_iteration <number>   Maximum number of iterations for EMG fitting. 
                                                             (default: '100') (min: '1')
  -PeptideQuantification:EMGScoring:init_mom <choice>        Alternative initial parameters for fitting throu
                                                             gh method of moments. (default: 'true') (valid: 
                                                             'true', 'false')

Parameters for FAIMS data processing:
  -PeptideQuantification:faims:merge_features <choice>       For FAIMS data with multiple compensation voltag
                                                             es: Merge features that represent the same analy
                                                             te detected at different CVs. Features are merge
                                                             d if they have the same charge and are within 5 
                                                             seconds RT and 0.05 Da m/z. Intensities are summ
                                                             ed. (default: 'true') (valid: 'true', 'false')

Alignment:
  -Alignment:model_type <choice>                             Options to control the modeling of retention 
                                                             time transformations from data (default: 'b_spli
                                                             ne') (valid: 'linear', 'b_spline', 'lowess', 
                                                             'interpolated')

Alignment:model:
  -Alignment:model:type <choice>                             Type of model (default: 'b_spline') (valid: 'lin
                                                             ear', 'b_spline', 'lowess', 'interpolated')

Parameters for 'linear' model:
  -Alignment:model:linear:symmetric_regression               Perform linear regression on 'y - x' vs. 'y + 
                                                             x', instead of on 'y' vs. 'x'.
  -Alignment:model:linear:x_weight <choice>                  Weight x values (default: 'x') (valid: '1/x', 
                                                             '1/x2', 'ln(x)', 'x')
  -Alignment:model:linear:y_weight <choice>                  Weight y values (default: 'y') (valid: '1/y', 
                                                             '1/y2', 'ln(y)', 'y')
  -Alignment:model:linear:x_datum_min <value>                Minimum x value (default: '1.0e-15')
  -Alignment:model:linear:x_datum_max <value>                Maximum x value (default: '1.0e15')
  -Alignment:model:linear:y_datum_min <value>                Minimum y value (default: '1.0e-15')
  -Alignment:model:linear:y_datum_max <value>                Maximum y value (default: '1.0e15')

Parameters for 'b_spline' model:
  -Alignment:model:b_spline:wavelength <value>               Determines the amount of smoothing by setting 
                                                             the number of nodes for the B-spline. The number
                                                              is chosen so that the spline approximates a 
                                                             low-pass filter with this cutoff wavelength. 
                                                             The wavelength is given in the same units as 
                                                             the data; a higher value means more smoothing. 
                                                             '0' sets the number of nodes to twice the number
                                                              of input points. (default: '0.0') (min: '0.0')
  -Alignment:model:b_spline:num_nodes <number>               Number of nodes for B-spline fitting. Overrides 
                                                             'wavelength' if set (to two or greater). A lower
                                                              value means more smoothing. (default: '5') (min
                                                             : '0')
  -Alignment:model:b_spline:extrapolate <choice>             Method to use for extrapolation beyond the origi
                                                             nal data range. 'linear': Linear extrapolation 
                                                             using the slope of the B-spline at the correspon
                                                             ding endpoint. 'b_spline': Use the B-spline (as 
                                                             for interpolation). 'constant': Use the constant
                                                              value of the B-spline at the corresponding endp
                                                             oint. 'global_linear': Use a linear fit through 
                                                             the data (which will most probably introduce 
                                                             discontinuities at the ends of the data range). 
                                                             (default: 'linear') (valid: 'linear', 'b_spline'
                                                             , 'constant', 'global_linear')
  -Alignment:model:b_spline:boundary_condition <number>      Boundary condition at B-spline endpoints: 0 (val
                                                             ue zero), 1 (first derivative zero) or 2 (second
                                                              derivative zero) (default: '2') (min: '0' max: 
                                                             '2')

Parameters for 'lowess' model:
  -Alignment:model:lowess:span <value>                       Fraction of datapoints (f) to use for each local
                                                              regression (determines the amount of smoothing)
                                                             . Choosing this parameter in the range .2 to .8 
                                                             usually results in a good fit. (default: '0.6666
                                                             66666666667') (min: '0.0' max: '1.0')
  -Alignment:model:lowess:auto_span                          If true, or if 'span' is 0, automatically select
                                                              LOWESS span by cross-validation.
  -Alignment:model:lowess:auto_span_min <value>              Lower bound for auto-selected span. (default: 
                                                             '0.15') (min: '1.0e-03')
  -Alignment:model:lowess:auto_span_max <value>              Upper bound for auto-selected span. (default: 
                                                             '0.8') (max: '0.99')
  -Alignment:model:lowess:auto_min_neighbors <number>        Minimum number of neighbors (span*n) enforced 
                                                             in auto mode. (default: '5') (min: '3')
  -Alignment:model:lowess:auto_k_folds <number>              K-folds for CV when n>50 (else LOO is used). 
                                                             (default: '5') (min: '2')
  -Alignment:model:lowess:auto_metric <choice>               Metric for CV selection: one of {'p90','p95','p9
                                                             9','rmse','mae'}. (default: 'mae') (valid: 'p90'
                                                             , 'p95', 'p99', 'rmse', 'mae')
  -Alignment:model:lowess:auto_span_grid <text>              Optional explicit grid of span candidates in 
                                                             (0,1]. Comma-separated list, e.g. '0.2,0.3,0.5'.
                                                               If empty, a default grid is used.
  -Alignment:model:lowess:num_iterations <number>            Number of robustifying iterations for lowess 
                                                             fitting. (default: '3') (min: '0')
  -Alignment:model:lowess:delta <value>                      Nonnegative parameter which may be used to save 
                                                             computations (recommended value is 0.01 of the 
                                                             range of the input, e.g. for data ranging from 
                                                             1000 seconds to 2000 seconds, it could be set 
                                                             to 10). Setting a negative value will automatica
                                                             lly do this. (default: '-1.0')
  -Alignment:model:lowess:interpolation_type <choice>        Method to use for interpolation between datapoin
                                                             ts computed by lowess. 'linear': Linear interpol
                                                             ation. 'cspline': Use the cubic spline for inter
                                                             polation. 'akima': Use an akima spline for inter
                                                             polation (default: 'cspline') (valid: 'linear', 
                                                             'cspline', 'akima')
  -Alignment:model:lowess:extrapolation_type <choice>        Method to use for extrapolation outside the data
                                                              range. 'two-point-linear': Uses a line through 
                                                             the first and last point to extrapolate. 'four-p
                                                             oint-linear': Uses a line through the first and 
                                                             second point to extrapolate in front and and a 
                                                             line through the last and second-to-last point 
                                                             in the end. 'global-linear': Uses a linear regre
                                                             ssion to fit a line through all data points and 
                                                             use it for interpolation. (default: 'four-point-
                                                             linear') (valid: 'two-point-linear', 'four-point
                                                             -linear', 'global-linear')

Parameters for 'interpolated' model:
  -Alignment:model:interpolated:interpolation_type <choice>  Type of interpolation to apply. (default: 'cspli
                                                             ne') (valid: 'linear', 'cspline', 'akima')
  -Alignment:model:interpolated:extrapolation_type <choice>  Type of extrapolation to apply: two-point-linear
                                                             : use the first and last data point to build a 
                                                             single linear model, four-point-linear: build 
                                                             two linear models on both ends using the first 
                                                             two / last two points, global-linear: use all 
                                                             points to build a single linear model. Note that
                                                              global-linear may not be continuous at the bord
                                                             er. (default: 'two-point-linear') (valid: 'two-p
                                                             oint-linear', 'four-point-linear', 'global-linea
                                                             r')

Alignment:align_algorithm:
  -Alignment:align_algorithm:score_type <text>               Name of the score type to use for ranking and 
                                                             filtering (.oms input only). If left empty, a 
                                                             score type is picked automatically.
  -Alignment:align_algorithm:min_run_occur <number>          Minimum number of runs (incl. reference, if any)
                                                              in which a peptide must occur to be used for 
                                                             the alignment.
                                                             Unless you have very few runs or identifications
                                                             , increase this value to focus on more informati
                                                             ve peptides. (default: '2') (min: '2')
  -Alignment:align_algorithm:max_rt_shift <value>            Maximum realistic RT difference for a peptide 
                                                             (median per run vs. reference). Peptides with 
                                                             higher shifts (outliers) are not used to compute
                                                              the alignment.
                                                             If 0, no limit (disable filter); if > 1, the 
                                                             final value in seconds; if <= 1, taken as a frac
                                                             tion of the range of the reference RT scale. 
                                                             (default: '0.1') (min: '0.0')
  -Alignment:align_algorithm:use_adducts <choice>            If IDs contain adducts, treat differently adduct
                                                             ed variants of the same molecule as different. 
                                                             (default: 'true') (valid: 'true', 'false')

Linking:
  -Linking:nr_partitions <number>                            How many partitions in m/z space should be used 
                                                             for the algorithm (more partitions means faster 
                                                             runtime and more memory efficient execution). 
                                                             (default: '100') (min: '1')
  -Linking:min_nr_diffs_per_bin <number>                     If IDs are used: How many differences from match
                                                             ing IDs should be used to calculate a linking 
                                                             tolerance for unIDed features in an RT region. 
                                                             RT regions will be extended until that number 
                                                             is reached. (default: '50') (min: '5')
  -Linking:min_IDscore_forTolCalc <value>                    If IDs are used: What is the minimum score of 
                                                             an ID to assume a reliable match for tolerance 
                                                             calculation. Check your current score type! (def
                                                             ault: '1.0')
  -Linking:noID_penalty <value>                              If IDs are used: For the normalized distances, 
                                                             how high should the penalty for missing IDs be? 
                                                             0 = no bias, 1 = IDs inside the max tolerances 
                                                             always preferred (even if much further away). 
                                                             (default: '0.0') (min: '0.0' max: '1.0')

Distance component based on m/z differences:
  -Linking:distance_MZ:max_difference <value>                Never pair features with larger m/z distance 
                                                             (unit defined by 'unit') (default: '10.0') (min:
                                                              '0.0')
  -Linking:distance_MZ:unit <choice>                         Unit of the 'max_difference' parameter (default:
                                                              'ppm') (valid: 'Da', 'ppm')

ProteinQuantification:
  -ProteinQuantification:method <choice>                     - top - quantify based on three most abundant 
                                                             peptides (number can be changed in 'top').
                                                             - iBAQ (intensity based absolute quantification)
                                                             , calculate the sum of all peptide peak intensit
                                                             ies divided by the number of theoretically obser
                                                             vable tryptic peptides (https://rdcu.be/cND1J). 
                                                             Warning: only consensusXML or featureXML input 
                                                             is allowed! (default: 'top') (valid: 'top', 'iBA
                                                             Q')
  -ProteinQuantification:best_charge                         Distinguish between fraction and charge states 
                                                             in detailed peptide output. For protein quantifi
                                                             cation, select one charge per modified peptide 
                                                             globally: maximize the number of (fraction group
                                                             , label) assays with a positive abundance, then 
                                                             break ties by total abundance; retain that charg
                                                             e's values in every assay.
                                                             By default, protein abundances are summed over 
                                                             ...
                                                             'fractions:aggregate', not by this flag.

Additional options for custom quantification using top N peptides.:
  -ProteinQuantification:top:N <number>                      Calculate protein abundance from this number of 
                                                             proteotypic peptides (most abundant first; '0' 
                                                             for all) (default: '3') (min: '0')
  -ProteinQuantification:top:aggregate <choice>              Aggregation method used to compute protein abund
                                                             ances from peptide abundances (default: 'median'
                                                             ) (valid: 'median', 'mean', 'weighted_mean', 
                                                             'sum')

Options for combining the fractions of a fraction group.:
  -ProteinQuantification:fractions:aggregate <choice>        How the fractions of one fraction group are comb
                                                             ined into that group's (fraction group, label) 
                                                             assay values.
                                                             - sum - add up every fraction, i.e. treat them 
                                                             as the parts of one separated sample that they 
                                                             are.
                                                             - best - keep a single fraction per peptide and 
                                                             fraction group and discard the others. The fract
                                                             ...
                                                             on and always report every file. (default: 'sum'
                                                             ) (valid: 'sum', 'best')

Additional options for consensus maps (and identification results comprising multiple runs):
  -ProteinQuantification:consensus:normalize                 Scale peptide abundances so that the median of 
                                                             each (fraction group, label) assay matches the 
                                                             overall median.
                                                             Abundances of zero count as 'not detected' and 
                                                             are left out of the medians; an assay without 
                                                             any positive abundance takes no part in the norm
                                                             alization.
  -ProteinQuantification:consensus:fix_peptides              Use the same peptides for protein quantification
                                                              across all (fraction group, label) assays.
                                                             With 'N 0',all peptides that occur in every assa
                                                             y are considered.
                                                             Otherwise ('N'), the N peptides that occur in 
                                                             the most assays (independently of each other) 
                                                             are selected,
                                                             breaking ties by total abundance (there is no 
                                                             ...
                                                             ), not a measurement of absence.

PipEcho:
  -PipEcho:fdr <value>                                       MBR FDR threshold (0.05=5%). (default: '0.05') 
                                                             (min: '0.0' max: '1.0')
  -PipEcho:random_seed <number>                              Seed for the random number generator used to 
                                                             select decoy donors. A fixed seed makes results 
                                                             reproducible. (default: '0')

PipEcho:distance_MZ:
  -PipEcho:distance_MZ:max_difference <value>                Never pair features with larger m/z distance 
                                                             (unit defined by 'unit') (default: '10.0') (min:
                                                              '0.0')
  -PipEcho:distance_MZ:unit <choice>                         Unit of the 'max_difference' parameter (default:
                                                              'ppm') (valid: 'Da', 'ppm')

                                                             
Common TOPP options:
  -ini <file>                                                Use the given TOPP INI file
  -threads <n>                                               Sets the number of threads allowed to be used 
                                                             by the TOPP tool (0 = all available cores) (defa
                                                             ult: '1')
  -write_ini <file>                                          Writes the default configuration file
  --help                                                     Shows options
  --helphelp                                                 Shows all options (including advanced)

INI file documentation of this tool:

Legend:
required parameter
advanced parameter

This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.

Each parameter entry contains the following information:

  • Name The identifier used in configuration files and on the command line.
  • Default value The value used if the parameter is not explicitly specified.
  • Description A short explanation describing the purpose and behavior of the parameter.
  • Tags Additional metadata associated with the parameter.
  • Restrictions Allowed value ranges for numeric parameters or valid options for string parameters.

Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.

Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.

+ProteomicsLFQA standard proteomics LFQ pipeline.
version3.6.0-pre-nightly-2026-09-29 Version of the tool that generated this parameters file.
++1Instance '1' section for 'ProteomicsLFQ'
in[] Input files. Optional only when '-feat_dir' supplies a checkpoint for every run of the design.input file*.mzML, *.d, *.raw
ids[] Identifications filtered at PSM level (e.g., q-value < 0.01).And annotated with PEP as main score.
We suggest using:
1. PSMFeatureExtractor to annotate percolator features.
2. PercolatorAdapter tool (score_type = 'q-value', -post_processing_tdc)
3. IDFilter (pep:score = 0.05)
To obtain well calibrated PEPs and an initial reduction of PSMs
ID files must be provided in same order as spectra files.
One identification per spectrum is expected: where several identifications of one
spectrum agree on peptidoform and charge, only the best-scoring one is kept, since
the others would count the same measurement more than once. Combine results from
several search engines with ConsensusID rather than concatenating them.
input file*.idXML, *.mzId, *.idparquet
design design fileinput file*.tsv
fasta fasta fileinput file*.fasta, *.fa, *.faa
out Optional output mzTab file. At least one output must be specified.output file*.mzTab
out_msstats Optional output MSstats input file. At least one output must be specified.output file*.csv
out_cxml Optional output consensusXML file. At least one output must be specified.output file*.consensusXML
out_qpx Optional output directory for QPX Parquet files (quantms.feature.parquet, quantms.psm.parquet, quantms.pg.parquet). At least one output must be specified.output dir
feat_dir Directory of per-run feature checkpoints. For every run of the experimental design, a valid checkpoint here is reused instead of detecting features again; a run without one is detected from '-in'/'-ids' and its checkpoint written. This makes a run resumable, and lets the per-run work be distributed: run with '-detect_only' on each machine, then once over the design with neither '-in' nor '-ids'. Requires '-design' and '-fasta'.output dir
detect_onlyfalse Stop after the per-run feature checkpoints have been written. No alignment, linking, inference or quantification is performed, and no result file is required. Requires '-feat_dir'.true, false
force_recomputefalse Ignore existing checkpoints in '-feat_dir' and detect every run again, overwriting them.true, false
proteinFDR0.05 Protein FDR threshold (0.05=5%).0.0:1.0
picked_proteinFDRfalse Use a picked protein FDR?true, false
psmFDR1.0 FDR threshold for sub-protein level (e.g. 0.05=5%). Use -FDR_type to choose the level. Cutoff is applied at the highest level. If Bayesian inference was chosen, it is equivalent with a peptide FDR0.0:1.0
FDR_typePSM Sub-protein FDR level. PSM, PSM+peptide (best PSM q-value).PSM, PSM+peptide
protein_inferenceaggregation Infer proteins:
aggregation = aggregates all peptide scores across a protein (using the best score)
bayesian = computes a posterior probability for every protein based on a Bayesian network.
Note: 'bayesian' only uses and reports the best PSM per peptide.
aggregation, bayesian
protein_quantificationunique_peptides Quantify proteins based on:
unique_peptides = use peptides mapping to single proteins or a group of indistinguishable proteins(according to the set of experimentally identified peptides).
strictly_unique_peptides = use peptides mapping to a unique single protein only.
shared_peptides = use shared peptides only for its best group (by inference score)
unique_peptides, strictly_unique_peptides, shared_peptides
quantification_methodfeature_intensity feature_intensity: MS1 signal.
spectral_counting: PSM counts.
feature_intensity, spectral_counting
targeted_onlyfalse true: Only ID based quantification.
false: include unidentified features so they can be linked to identified ones (=match between runs).
true, false
pip_echofalse Perform match between runs (MBR) via PIP-ECHOtrue, false
feature_with_id_min_score0.0 The minimum probability (e.g.: 0.25) an identified (=id targeted) feature must have to be kept for alignment and linking (0=no filter).0.0:1.0
feature_without_id_min_score0.0 The minimum probability (e.g.: 0.75) an unidentified feature must have to be kept for alignment and linking (0=no filter).0.0:1.0
mass_recalibrationfalse Mass recalibration.true, false
alignment_orderstar If star, aligns all maps to the map that shares the most IDs with every other map. If tree_guided, aligns maps in tree order (most similar pairs first).star, tree_guided
keep_feature_top_psm_onlytrue If false, also keeps lower ranked PSMs that have the top-scoring sequence as a candidate per feature in the same file.true, false
log Name of log file (created only when specified)
debug0 Sets the debug level
threads1 Sets the number of threads allowed to be used by the TOPP tool (0 = all available cores)
no_progressfalse Disables progress logging to command linetrue, false
forcefalse Overrides tool-specific checkstrue, false
testfalse Enables the test mode (needed for internal use only)true, false
+++SeedingParameters for seeding of untargeted features
intThreshold1.0e04 Peak intensity threshold applied in seed detection.
charge2:5 Charge range considered for untargeted feature seeds.
traceRTTolerance3.0 Combines all spectra in the tolerance window to stabilize identification of isotope patterns. Controls sensitivity (low value) vs. specificity (high value) of feature seeds.
algorithmmultiplex Algorithm for untargeted seed feature detection.
multiplex: FeatureFinderMultiplexAlgorithm (default, current behavior).
biosaur2: Biosaur2Algorithm (handles IM_PEAK/PASEF data natively).
multiplex, biosaur2
++++Biosaur2
mini500.0 Minimum intensity threshold0.0:∞
minmz350.0 Minimum m/z value0.0:∞
maxmz1500.0 Maximum m/z value0.0:∞
htol8.0 Mass accuracy in ppm for combining peaks into hills0.0:∞
itol8.0 Mass accuracy in ppm for isotopic patterns0.0:∞
hvf1.3 Hill valley factor for splitting hills1.0:∞
ivf5.0 Isotope valley factor for splitting isotope patterns1.0:∞
minlh3 Minimum number of scans for a hill1:∞
pasefmini100.0 Minimum combined intensity for PASEF/TIMS clusters after m/z–ion-mobility centroiding.0.0:∞
pasefminlh2 Minimum number of raw points per PASEF/TIMS cluster during centroiding.1:∞
cmin1 Minimum charge state1:∞
cmax6 Maximum charge state1:∞
iuse0 Number of isotopes for intensity calculation (0=mono only, -1=all, 1=mono+first, etc.)-1:∞
nmfalse Negative mode (affects neutral mass calculation)true, false
toffalse Enable TOF-specific intensity filteringtrue, false
profilefalse Enable profile mode processing (centroid spectra using PeakPickerHiRes)true, false
paseftol0.05 Ion mobility accuracy for linking peaks into hills and grouping isotopes (0 = disable IM-based gating). Default is for 1/K0 units (e.g., 0.03-0.08). For CCS data, use larger values (e.g., 5-20 square angstroms).0.0:∞
use_hill_calibfalse Enable automatic hill mass tolerance calibrationtrue, false
ignore_iso_calibfalse Disable automatic isotope mass error calibrationtrue, false
hrttol10.0 Maximum allowed RT difference (in seconds) between monoisotopic hill apex and isotope hill apex when assembling isotope patterns (0 disables RT gating).0.0:∞
convex_hullsbounding_box Representation of feature convex hulls in the output FeatureMap. 'bounding_box' stores a single RT–m/z bounding box per feature (smaller featureXML, no per-trace detail), whereas 'mass_traces' stores one convex hull per contributing hill using all mass-trace points (larger featureXML, preserves detailed trace shape).mass_traces, bounding_box
faims_merge_featurestrue For FAIMS data with multiple compensation voltages: Merge features representing the same analyte detected at different CV values into a single feature. Only features with DIFFERENT FAIMS CV values are merged (same CV = different analytes). Has no effect on non-FAIMS data.true, false
+++Centroiding
signal_to_noise0.0 Minimal signal-to-noise ratio for a peak to be picked (0.0 disables SNT estimation!)0.0:∞
spacing_difference_gap4.0 The extension of a peak is stopped if the spacing between two subsequent data points exceeds 'spacing_difference_gap * min_spacing'. 'min_spacing' is the smaller of the two spacings from the peak apex to its two neighboring points. '0' to disable the constraint. Not applicable to chromatograms.0.0:∞
spacing_difference1.5 Maximum allowed difference between points during peak extension, in multiples of the minimal difference between the peak apex and its two neighboring points. If this difference is exceeded a missing point is assumed (see parameter 'missing'). A higher value implies a less stringent peak definition, since individual signals within the peak are allowed to be further apart. '0' to disable the constraint. Not applicable to chromatograms.0.0:∞
missing1 Maximum number of missing points allowed when extending a peak to the left or to the right. A missing data point occurs if the spacing between two subsequent data points exceeds 'spacing_difference * min_spacing'. 'min_spacing' is the smaller of the two spacings from the peak apex to its two neighboring points. Not applicable to chromatograms.0:∞
ms_levels[] List of MS levels for which the peak picking is applied. If empty, auto mode is enabled, all peaks which aren't picked yet will get picked. Other scans are copied to the output without changes.1:∞
report_FWHMfalse Add metadata for FWHM (as floatDataArray named 'FWHM' or 'FWHM_ppm', depending on param 'report_FWHM_unit') for each picked peak.true, false
report_FWHM_unitrelative Unit of FWHM. Either absolute in the unit of input, e.g. 'm/z' for spectra, or relative as ppm (only sensible for spectra, not chromatograms).relative, absolute
allow_missing_flankfalse Allow peaks without flanking data points on both sides. This is useful for TimsTOF data where profile peaks may be missing the leading or trailing edge.true, false
++++SignalToNoise
max_intensity-1 maximal intensity considered for histogram construction. By default, it will be calculated automatically (see auto_mode). Only provide this parameter if you know what you are doing (and change 'auto_mode' to '-1')! All intensities EQUAL/ABOVE 'max_intensity' will be added to the LAST histogram bin. If you choose 'max_intensity' too small, the noise estimate might be too small as well. If chosen too big, the bins become quite large (which you could counter by increasing 'bin_count', which increases runtime). In general, the Median-S/N estimator is more robust to a manual max_intensity than the MeanIterative-S/N.-1:∞
auto_max_stdev_factor3.0 parameter for 'max_intensity' estimation (if 'auto_mode' == 0): mean + 'auto_max_stdev_factor' * stdev0.0:999.0
auto_max_percentile95 parameter for 'max_intensity' estimation (if 'auto_mode' == 1): auto_max_percentile th percentile0:100
auto_mode0 method to use to determine maximal intensity: -1 --> use 'max_intensity'; 0 --> 'auto_max_stdev_factor' method (default); 1 --> 'auto_max_percentile' method-1:1
win_len200.0 window length in Thomson1.0:∞
bin_count30 number of bins for intensity values3:∞
min_required_elements10 minimum number of elements required in a window (otherwise it is considered sparse)1:∞
noise_for_empty_window1.0e20 noise value used for sparse windows
write_log_messagestrue Write out log messages in case of sparse windows or median in rightmost histogram bintrue, false
+++PeptideQuantification
candidates_out Optional output file with feature candidates.output file
debug0 Debug level for feature detection.0:∞
quantify_decoysfalse Whether decoy peptides should be quantified (true) or skipped (false).true, false
min_psm_cutoffnone Minimum score for the best PSM of a spectrum to be used as seed. Use 'none' for no cutoff.
add_mass_offset_peptides0.0 If for every peptide (or seed) also an offset peptide is extracted (true). Can be used to downstream to determine MBR false transfer rates. (0.0 = disabled)0.0:∞
seed_apex_rt_tolerance5.0 Maximum allowed RT deviation (in seconds) between a seed's apex and the detected feature's apex. Seed-derived features whose detected apex deviates more than this value from the original seed apex are removed during filtering. Useful to discard unreliable seed extractions where the picked peak is far from the seed location. (0 = disabled)0.0:∞
++++extractParameters for ion chromatogram extraction
batch_size5000 Nr of peptides used in each batch of chromatogram extraction. Smaller values decrease memory usage but increase runtime.1:∞
mz_window10.0 m/z window size for chromatogram extraction (unit: ppm if 1 or greater, else Da/Th)0.0:∞
IM_window0.06 Ion mobility (IM) window for chromatogram extraction in the IM dimension. Set to 0.0 to disable IM filtering (even if data contains IM information). The window is applied as +/- IM_window/2 around the median IM value of identified peptides. This parameter is automatically ignored if the input data does not contain IM information (determined via IMTypes::determineIMFormat). Currently only concatenated IM format is supported. Typical values: 0.05-0.10 for TIMS data (1/K0 units), 10-50 for CCS data (square angstroms), 3-5 for FAIMS data (compensation voltage).Note: IM values are calculated per peptide/charge/RT-region, using the median of all identifications in that region for robustness. The median, min, and max IM values are propagated to output features as meta-values (IM_median, IM_min, IM_max) for quality control.0.0:∞
n_isotopes2 Number of isotopes to include in each peptide assay.2:∞
isotope_pmin0.0 Minimum probability for an isotope to be included in the assay for a peptide. If set, this parameter takes precedence over 'extract:n_isotopes'.0.0:1.0
rt_window0.0 RT window size (in sec.) for chromatogram extraction. If not set, it is derived from 'detect:peak_width' and 'detect:mapping_tolerance'.0.0:∞
++++detectParameters for detecting features in extracted ion chromatograms
min_peak_width0.2 Minimum elution peak width. Absolute value in seconds if 1 or greater, else relative to 'peak_width'.0.0:∞
signal_to_noise0.8 Signal-to-noise threshold for OpenSWATH feature detection0.1:∞
mapping_tolerance0.0 RT tolerance (plus/minus) for mapping peptide IDs to features. Absolute value in seconds if 1 or greater, else relative to the RT span of the feature.0.0:∞
++++modelParameters for fitting elution models to features
typesymmetric Type of elution model to fit to featuressymmetric, asymmetric, none
add_zeros0.2 Add zero-intensity points outside the feature range to constrain the model fit. This parameter sets the weight given to these points during model fitting; '0' to disable.0.0:∞
unweighted_fitfalse Suppress weighting of mass traces according to theoretical intensities when fitting elution modelstrue, false
no_imputationfalse If fitting the elution model fails for a feature, set its intensity to zero instead of imputing a value from the initial intensity estimatetrue, false
each_tracefalse Fit elution model to each individual mass tracetrue, false
+++++checkParameters for checking the validity of elution models (and rejecting them if necessary)
min_area1.0 Lower bound for the area under the curve of a valid elution model0.0:∞
boundaries0.5 Time points corresponding to this fraction of the elution model height have to be within the data region used for model fitting0.0:1.0
width10.0 Upper limit for acceptable widths of elution models (Gaussian or EGH), expressed in terms of modified (median-based) z-scores. '0' to disable. Not applied to individual mass traces (parameter 'each_trace').0.0:∞
asymmetry10.0 Upper limit for acceptable asymmetry of elution models (EGH only), expressed in terms of modified (median-based) z-scores. '0' to disable. Not applied to individual mass traces (parameter 'each_trace').0.0:∞
++++EMGScoringParameters for fitting exp. mod. Gaussians to mass traces.
max_iteration100 Maximum number of iterations for EMG fitting.1:∞
init_momtrue Alternative initial parameters for fitting through method of moments.true, false
++++faimsParameters for FAIMS data processing
merge_featurestrue For FAIMS data with multiple compensation voltages: Merge features that represent the same analyte detected at different CVs. Features are merged if they have the same charge and are within 5 seconds RT and 0.05 Da m/z. Intensities are summed.true, false
+++Alignment
model_typeb_spline Options to control the modeling of retention time transformations from datalinear, b_spline, lowess, interpolated
++++model
typeb_spline Type of modellinear, b_spline, lowess, interpolated
+++++linearParameters for 'linear' model
symmetric_regressionfalse Perform linear regression on 'y - x' vs. 'y + x', instead of on 'y' vs. 'x'.true, false
x_weightx Weight x values1/x, 1/x2, ln(x), x
y_weighty Weight y values1/y, 1/y2, ln(y), y
x_datum_min1.0e-15 Minimum x value
x_datum_max1.0e15 Maximum x value
y_datum_min1.0e-15 Minimum y value
y_datum_max1.0e15 Maximum y value
+++++b_splineParameters for 'b_spline' model
wavelength0.0 Determines the amount of smoothing by setting the number of nodes for the B-spline. The number is chosen so that the spline approximates a low-pass filter with this cutoff wavelength. The wavelength is given in the same units as the data; a higher value means more smoothing. '0' sets the number of nodes to twice the number of input points.0.0:∞
num_nodes5 Number of nodes for B-spline fitting. Overrides 'wavelength' if set (to two or greater). A lower value means more smoothing.0:∞
extrapolatelinear Method to use for extrapolation beyond the original data range. 'linear': Linear extrapolation using the slope of the B-spline at the corresponding endpoint. 'b_spline': Use the B-spline (as for interpolation). 'constant': Use the constant value of the B-spline at the corresponding endpoint. 'global_linear': Use a linear fit through the data (which will most probably introduce discontinuities at the ends of the data range).linear, b_spline, constant, global_linear
boundary_condition2 Boundary condition at B-spline endpoints: 0 (value zero), 1 (first derivative zero) or 2 (second derivative zero)0:2
+++++lowessParameters for 'lowess' model
span0.666666666666667 Fraction of datapoints (f) to use for each local regression (determines the amount of smoothing). Choosing this parameter in the range .2 to .8 usually results in a good fit.0.0:1.0
auto_spanfalse If true, or if 'span' is 0, automatically select LOWESS span by cross-validation.true, false
auto_span_min0.15 Lower bound for auto-selected span.1.0e-03:∞
auto_span_max0.8 Upper bound for auto-selected span.-∞:0.99
auto_min_neighbors5 Minimum number of neighbors (span*n) enforced in auto mode.3:∞
auto_k_folds5 K-folds for CV when n>50 (else LOO is used).2:∞
auto_metricmae Metric for CV selection: one of {'p90','p95','p99','rmse','mae'}.p90, p95, p99, rmse, mae
auto_span_grid Optional explicit grid of span candidates in (0,1]. Comma-separated list, e.g. '0.2,0.3,0.5'. If empty, a default grid is used.
num_iterations3 Number of robustifying iterations for lowess fitting.0:∞
delta-1.0 Nonnegative parameter which may be used to save computations (recommended value is 0.01 of the range of the input, e.g. for data ranging from 1000 seconds to 2000 seconds, it could be set to 10). Setting a negative value will automatically do this.
interpolation_typecspline Method to use for interpolation between datapoints computed by lowess. 'linear': Linear interpolation. 'cspline': Use the cubic spline for interpolation. 'akima': Use an akima spline for interpolationlinear, cspline, akima
extrapolation_typefour-point-linear Method to use for extrapolation outside the data range. 'two-point-linear': Uses a line through the first and last point to extrapolate. 'four-point-linear': Uses a line through the first and second point to extrapolate in front and and a line through the last and second-to-last point in the end. 'global-linear': Uses a linear regression to fit a line through all data points and use it for interpolation.two-point-linear, four-point-linear, global-linear
+++++interpolatedParameters for 'interpolated' model
interpolation_typecspline Type of interpolation to apply.linear, cspline, akima
extrapolation_typetwo-point-linear Type of extrapolation to apply: two-point-linear: use the first and last data point to build a single linear model, four-point-linear: build two linear models on both ends using the first two / last two points, global-linear: use all points to build a single linear model. Note that global-linear may not be continuous at the border.two-point-linear, four-point-linear, global-linear
++++align_algorithm
score_type Name of the score type to use for ranking and filtering (.oms input only). If left empty, a score type is picked automatically.
score_cutofffalse Use only IDs above a score cut-off (parameter 'min_score') for alignment?true, false
min_score0.05 If 'score_cutoff' is 'true': Minimum score for an ID to be considered.
Unless you have very few runs or identifications, increase this value to focus on more informative peptides.
min_run_occur2 Minimum number of runs (incl. reference, if any) in which a peptide must occur to be used for the alignment.
Unless you have very few runs or identifications, increase this value to focus on more informative peptides.
2:∞
max_rt_shift0.1 Maximum realistic RT difference for a peptide (median per run vs. reference). Peptides with higher shifts (outliers) are not used to compute the alignment.
If 0, no limit (disable filter); if > 1, the final value in seconds; if <= 1, taken as a fraction of the range of the reference RT scale.
0.0:∞
use_unassigned_peptidesfalse Should unassigned peptide identifications be used when computing an alignment of feature or consensus maps? If 'false', only peptide IDs assigned to features will be used.true, false
use_feature_rttrue When aligning feature or consensus maps, don't use the retention time of a peptide identification directly; instead, use the retention time of the centroid of the feature (apex of the elution profile) that the peptide was matched to. If different identifications are matched to one feature, only the peptide closest to the centroid in RT is used.
Precludes 'use_unassigned_peptides'.
true, false
use_adductstrue If IDs contain adducts, treat differently adducted variants of the same molecule as different.true, false
auto_referencebest_run Reference to align to if none is given (neither a reference file nor an input index): 'best_run' - the input that shares the most identified sequences with every other input (on ties, the one with the most identified sequences). A consensus is used instead if no input shares at least two sequences with every other input. If the chosen input leaves other inputs with too few alignment points, other inputs and a consensus are tried as well (see 'auto_reference_min_points'). 'consensus' - median RTs per sequence over all inputs. A consensus favors none of the inputs, but only partly corrects larger RT shifts, because every input contributes to the consensus it is aligned to.best_run, consensus
auto_reference_min_points11 If 'auto_reference' is 'best_run': number of alignment points (after removing outliers, see 'max_rt_shift') that the reference should provide for every other input. If the chosen input leaves inputs with fewer points, and one of them shares at least this many sequences with other inputs, every input and a consensus of all inputs are tried as the reference. The choice that gives the most inputs at least this many points is used (the reference counts); on ties, the first choice is kept, and an input is preferred over a consensus. The default is the smallest number of points to which ProteomicsLFQ and MS1LabeledWorkflow fit an RT model. 0 disables the check.0:∞
+++Linking
use_identificationstrue Never link features that are annotated with different peptides (only the best hit per peptide identification is taken into account).true, false
nr_partitions100 How many partitions in m/z space should be used for the algorithm (more partitions means faster runtime and more memory efficient execution).1:∞
min_nr_diffs_per_bin50 If IDs are used: How many differences from matching IDs should be used to calculate a linking tolerance for unIDed features in an RT region. RT regions will be extended until that number is reached.5:∞
min_IDscore_forTolCalc1.0 If IDs are used: What is the minimum score of an ID to assume a reliable match for tolerance calculation. Check your current score type!
noID_penalty0.0 If IDs are used: For the normalized distances, how high should the penalty for missing IDs be? 0 = no bias, 1 = IDs inside the max tolerances always preferred (even if much further away).0.0:1.0
ignore_chargefalse false [default]: pairing requires equal charge state (or at least one unknown charge '0'); true: Pairing irrespective of charge statetrue, false
ignore_adducttrue true [default]: pairing requires equal adducts (or at least one without adduct annotation); true: Pairing irrespective of adductstrue, false
++++distance_RTDistance component based on RT differences
exponent1.0 Normalized RT differences ([0-1], relative to 'max_difference') are raised to this power (using 1 or 2 will be fast, everything else is REALLY slow)0.0:∞
weight1.0 Final RT distances are weighted by this factor0.0:∞
++++distance_MZDistance component based on m/z differences
max_difference10.0 Never pair features with larger m/z distance (unit defined by 'unit')0.0:∞
unitppm Unit of the 'max_difference' parameterDa, ppm
exponent2.0 Normalized ([0-1], relative to 'max_difference') m/z differences are raised to this power (using 1 or 2 will be fast, everything else is REALLY slow)0.0:∞
weight5.0 Final m/z distances are weighted by this factor0.0:∞
++++distance_intensityDistance component based on differences in relative intensity (usually relative to highest peak in the whole data set)
exponent1.0 Differences in relative intensity ([0-1]) are raised to this power (using 1 or 2 will be fast, everything else is REALLY slow)0.0:∞
weight0.1 Final intensity distances are weighted by this factor0.0:∞
log_transformdisabled Log-transform intensities? If disabled, d = |int_f2 - int_f1| / int_max. If enabled, d = |log(int_f2 + 1) - log(int_f1 + 1)| / log(int_max + 1))enabled, disabled
+++ProteinQuantification
methodtop - top - quantify based on three most abundant peptides (number can be changed in 'top').
- iBAQ (intensity based absolute quantification), calculate the sum of all peptide peak intensities divided by the number of theoretically observable tryptic peptides (https://rdcu.be/cND1J). Warning: only consensusXML or featureXML input is allowed!
top, iBAQ
best_chargefalse Distinguish between fraction and charge states in detailed peptide output. For protein quantification, select one charge per modified peptide globally: maximize the number of (fraction group, label) assays with a positive abundance, then break ties by total abundance; retain that charge's values in every assay.
By default, protein abundances are summed over all charge states. How the retained values of several fractions are combined is governed by 'fractions:aggregate', not by this flag.
true, false
++++topAdditional options for custom quantification using top N peptides.
N3 Calculate protein abundance from this number of proteotypic peptides (most abundant first; '0' for all)0:∞
aggregatemedian Aggregation method used to compute protein abundances from peptide abundancesmedian, mean, weighted_mean, sum
include_alltrue Include results for proteins with fewer proteotypic peptides than indicated by 'N' (no effect if 'N' is 0 or 1)true, false
++++fractionsOptions for combining the fractions of a fraction group.
aggregatesum How the fractions of one fraction group are combined into that group's (fraction group, label) assay values.
- sum - add up every fraction, i.e. treat them as the parts of one separated sample that they are.
- best - keep a single fraction per peptide and fraction group and discard the others. The fraction is chosen ONCE per peptide, ranked by the number of labels in which it has a positive abundance and then by the total of those abundances (an exact tie keeps the lowest fraction number), and ALL of its labels are then taken from it. The choice is deliberately not made per label: taking one channel from one fraction and another channel from a different fraction would mix physical aliquots and destroy the reporter-ion ratios that isobaric quantification consists of.
Only the assay values are affected. Per-(file, channel) quantities are per fraction by definition and always report every file.
sum, best
++++consensusAdditional options for consensus maps (and identification results comprising multiple runs)
normalizefalse Scale peptide abundances so that the median of each (fraction group, label) assay matches the overall median.
Abundances of zero count as 'not detected' and are left out of the medians; an assay without any positive abundance takes no part in the normalization.
true, false
fix_peptidesfalse Use the same peptides for protein quantification across all (fraction group, label) assays.
With 'N 0',all peptides that occur in every assay are considered.
Otherwise ('N'), the N peptides that occur in the most assays (independently of each other) are selected,
breaking ties by total abundance (there is no guarantee that the best co-ocurring peptides are chosen!).
A peptide counts as occurring in an assay only where its abundance is positive: an abundance stored as zero means 'not detected' (e.g. an isobaric reporter below 'min_reporter_intensity'), not a measurement of absence.
true, false
+++PipEcho
fdr0.05 MBR FDR threshold (0.05=5%).0.0:1.0
random_seed0 Seed for the random number generator used to select decoy donors. A fixed seed makes results reproducible.
min_decoys20 Minimum number of MBR decoy transfers required before any transferred feature is kept. Transferred features are also dropped whenever the requested 'fdr' cannot be resolved by the available decoys (1/decoys > fdr). When transfers are dropped, only direct identifications are retained. A 'fdr' of 1.0 disables FDR control and keeps all transfers regardless of this value.1:∞
max_training_points50000 Upper bound on the number of candidate transfers used to TRAIN the transfer-FDR SVM in each cross-validation fold (0 = unlimited). Predictions and the FDR/q-values are always computed over ALL transfers, so this only bounds the SVM-fitting cost and does not change which transfers are scored. Runs with very many features per run can otherwise make the SVM grid search slow. When the cap is hit, a deterministic stratified subsample is used: both labelled classes (decoys and positives) are kept whole if they fit, otherwise the minority class is kept whole and the majority is score-spread sampled to fit the budget.0:∞
++++distance_MZ
max_difference10.0 Never pair features with larger m/z distance (unit defined by 'unit')0.0:∞
unitppm Unit of the 'max_difference' parameterDa, ppm
++++local_rt
enabledtrue Use a LOCAL adaptive retention-time window for MBR candidate search instead of the single global RT window ('distance_RT:max_difference'). For each donor the expected acceptor RT is predicted from nearby peptides identified in BOTH runs (local alignment) and the search window is sized from the local RT scatter, sharpening the RT feature and removing much of the false in-window background. The window is widened adaptively if too few decoys are produced to resolve the FDR. Set 'false' to restore the legacy single global RT window.true, false
max_window100.0 Maximum half-width (seconds) of the local RT window before adaptive widening (backstop cap on the data-driven width).0.0:∞
min_window5.0 Minimum half-width (seconds) of the local RT window.0.0:∞
anchor_window120.0 Only peptides within this RT distance (seconds) of the donor are used as local-alignment anchors.0.0:∞
anchors3 Maximum number of anchor peptides per side used for the local RT prediction.1:∞
sigma_scale3.0 Local RT window half-width = this multiple of the local anchor RT-shift standard deviation.0.0:∞
fallback_window15.0 Half-width (seconds) used when fewer than two local anchors are available.0.0:∞
autotrue Estimate the local RT window scales (max_window, min_window, anchor_window, fallback_window) from the data instead of using the fixed values above, so one configuration adapts across short and long gradients. Derived from the RT-shift scatter of shared anchors and the anchor density; uses 'median_fwhm' as a physical lower guard when provided.true, false
median_fwhm0.0 Chromatographic peak FWHM (seconds) used as a physical lower guard when 'auto' is enabled (0 = estimate window scales from RT residuals and anchor density only). Set by the host tool (e.g. ProteomicsLFQ) which has raw-spectra context.0.0:∞
rt_scoresvm_and_mbr How the retention-time feature is used on the local RT path. 'raw': the raw |Δrt| is the SVM predictor (legacy). 'svm': a calibrated [0,1] RT-agreement score (FlashLFQ-style two-tailed CDF against the data-driven RT prediction-error distribution) replaces it as the SVM predictor. 'svm_and_mbr': as 'svm', and the calibrated score also enters the bootstrap geometric mean. Ignored unless 'enabled'.raw, svm, svm_and_mbr