![]() |
OpenMS
|
Performs label-free quantification of peptides and proteins.
Input:
Identifications in idXML or mzIdentML format with posterior error probabilities as score type. To generate those we suggest to run:
Exactly one identification run per ID file is required, and merged ID runs are not supported. One identification per spectrum is expected as well: ProteomicsLFQ measures one value per (spectrum, peptidoform, charge), so where several identifications of one spectrum agree on all three, only the best-scoring one is kept and the reduction is reported. The others would otherwise count the same measurement more than once, in the PSM-level FDR and in every output. Results from several search engines must therefore be combined - with ConsensusID (-algorithm best -keep_old_scores, which preserves each engine's score) - rather than simply concatenated. Identifications of one spectrum that name different peptidoforms are left alone: a chimeric spectrum yields two distinct measurements.
Match between runs with PIP-ECHO (-pip_echo true):
By default, the second method links the features across runs by QT clustering on their RT and m/z (Linking:*), without estimating the error rate of the transferred identifications. With -pip_echo true, PIP-ECHO links the features instead and controls the false discovery rate of the transfers. Decoy transfers, which look for a peptide at the retention time of an unrelated peptide, estimate the error rate. A support vector machine scores each transfer (on intensity, mass error, RT agreement, isotope envelope and, if every run has ion mobility data, ion mobility), and the transfers up to PipEcho:fdr (default 0.05) are kept. -pip_echo requires -targeted_only false (the default).
PipEcho:local_rt:*), and it is widened if there are too few decoys to resolve the FDR. PipEcho:local_rt:enabled false uses one global window instead, which ProteomicsLFQ sizes from the alignment error and the chromatographic peak width.PipEcho:min_decoys (default 20), or too few to resolve the requested FDR, no transfer is kept, and only the direct identifications remain. PipEcho:fdr 1.0 keeps all transfers.PipEcho:random_seed selects the decoys; with the same seed (default 0), results are reproducible.PipEcho:max_training_points limits how many transfers the support vector machine is trained on in each cross-validation fold (default 50000, 0 = no limit), which bounds the run time on large data sets. All transfers are scored.Resuming and distributing feature detection (-feat_dir):
Feature detection is the expensive part of the workflow and each MS run is detected independently of every other; alignment, linking, inference and quantification need all runs at once. -feat_dir names a directory of per-run feature checkpoints and applies one rule to every row of the experimental design: reuse its checkpoint if a valid one is there, otherwise detect the run from -in / -ids and write one, otherwise fail.
Everything follows from that rule:
A checkpoint records the parameters, OpenMS build, experimental-design row and input files it was produced from, and is refused if any of those disagree with the run trying to use it - naming the setting that differs. This is what makes reuse safe rather than merely convenient: nothing else would stop half a study being detected with one setting and half with another. There is no way to combine checkpoints that disagree: -force_recompute detects the affected runs again and rewrites their checkpoints.
-feat_dir requires an explicit -design (a generated one would label every separately detected run as the first) and -fasta (a checkpoint has to carry the peptide-indexing results, which a combining run cannot reconstruct), and does not apply to spectral_counting.
Note on scale: the combining step holds every run's features of a fraction in memory at once, so its ceiling is set by features per run rather than by run count. Measured on a 7-run dataset averaging about 10,000 features per run, the marginal cost is roughly 2 kB per feature plus 10 MB per run, so 200 such runs in one fraction need about 9 GB; a deep-proteome experiment at ~60,000 features per run would need several times that.
FAIMS (Field Asymmetric Ion Mobility Spectrometry):
FAIMS data is automatically detected based on compensation voltage (CV) annotations in the mzML file. The data is split by CV and processed separately for each voltage group during feature detection. Features representing the same analyte detected at different CV values are merged automatically. The merged features are then aligned and linked across runs based on RT and m/z. No special preparation of the input mzML file is required.
Bruker .d (TimsTOF PASEF):
Bruker .d directories containing DDA-PASEF data are supported directly. When .d input is detected, the tool automatically:
Normalization:
Output (at least one required; each output is optional individually):
out)out_msstats)out_cxml)out_qpx)The command line parameters of this tool are:
ProteomicsLFQ -- A standard proteomics LFQ pipeline.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ProteomicsLFQ.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
+ Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.
Usage:
ProteomicsLFQ <options>
Options (mandatory options marked with '*'):
-in <file list> Input files. Optional only when '-feat_dir' supp
lies a checkpoint for every run of the design.
(valid formats: 'mzML', 'd', 'raw')
-ids <file list> Identifications filtered at PSM level (e.g.,
q-value < 0.01).And annotated with PEP as main
score.
We suggest using:
1. PSMFeatureExtractor to annotate percolator
features.
2. PercolatorAdapter tool (score_type = 'q-value
', -post_processing_tdc)
...
than concatenating them. (valid formats: 'idXML'
, 'mzId', 'idparquet')
-design <file> Design file (valid formats: 'tsv')
-fasta <file> Fasta file (valid formats: 'fasta', 'fa', 'faa')
-out <file> Optional output mzTab file. At least one output
must be specified. (valid formats: 'mzTab')
-out_msstats <file> Optional output MSstats input file. At least
one output must be specified. (valid formats:
'csv')
-out_cxml <file> Optional output consensusXML file. At least one
output must be specified. (valid formats: 'conse
nsusXML')
-out_qpx <directory> Optional output directory for QPX Parquet files
(quantms.feature.parquet, quantms.psm.parquet,
quantms.pg.parquet). At least one output must
be specified.
-feat_dir <directory> Directory of per-run feature checkpoints. For
every run of the experimental design, a valid
checkpoint here is reused instead of detecting
features again; a run without one is detected
from '-in'/'-ids' and its checkpoint written.
This makes a run resumable, and lets the per-run
work be distributed: run with '-detect_only'
on each machine, then once over the design with
neither '-in' nor '-ids'. Requires '-design'
and '-fasta'.
-detect_only Stop after the per-run feature checkpoints have
been written. No alignment, linking, inference
or quantification is performed, and no result
file is required. Requires '-feat_dir'.
-proteinFDR <threshold> Protein FDR threshold (0.05=5%). (default: '0.05
') (min: '0.0' max: '1.0')
-picked_proteinFDR <choice> Use a picked protein FDR? (default: 'false')
(valid: 'true', 'false')
-psmFDR <threshold> FDR threshold for sub-protein level (e.g. 0.05=5
%). Use -FDR_type to choose the level. Cutoff
is applied at the highest level. If Bayesian
inference was chosen, it is equivalent with a
peptide FDR (default: '1.0') (min: '0.0' max:
'1.0')
-FDR_type <threshold> Sub-protein FDR level. PSM, PSM+peptide (best
PSM q-value). (default: 'PSM') (valid: 'PSM',
'PSM+peptide')
-quantification_method <option> Feature_intensity: MS1 signal.
spectral_counting: PSM counts. (default: 'featur
e_intensity') (valid: 'feature_intensity', 'spec
tral_counting')
-targeted_only <option> True: Only ID based quantification.
false: include unidentified features so they
can be linked to identified ones (=match between
runs). (default: 'false') (valid: 'true', 'fals
e')
-pip_echo <option> Perform match between runs (MBR) via PIP-ECHO
(default: 'false') (valid: 'true', 'false')
Parameters for seeding of untargeted features:
-Seeding:algorithm <choice> Algorithm for untargeted seed feature detection.
multiplex: FeatureFinderMultiplexAlgorithm (defa
ult, current behavior).
biosaur2: Biosaur2Algorithm (handles IM_PEAK/PAS
EF data natively). (default: 'multiplex') (valid
: 'multiplex', 'biosaur2')
Centroiding:
-Centroiding:signal_to_noise <value> Minimal signal-to-noise ratio for a peak to be
picked (0.0 disables SNT estimation!) (default:
'0.0') (min: '0.0')
-Centroiding:ms_levels <numbers> List of MS levels for which the peak picking is
applied. If empty, auto mode is enabled, all
peaks which aren't picked yet will get picked.
Other scans are copied to the output without
changes. (min: '1')
PeptideQuantification:
-PeptideQuantification:quantify_decoys Whether decoy peptides should be quantified (tru
e) or skipped (false).
-PeptideQuantification:min_psm_cutoff <text> Minimum score for the best PSM of a spectrum to
be used as seed. Use 'none' for no cutoff. (defa
ult: 'none')
-PeptideQuantification:add_mass_offset_peptides <value> If for every peptide (or seed) also an offset
peptide is extracted (true). Can be used to down
stream to determine MBR false transfer rates.
(0.0 = disabled) (default: '0.0') (min: '0.0')
Parameters for ion chromatogram extraction:
-PeptideQuantification:extract:batch_size <number> Nr of peptides used in each batch of chromatogra
m extraction. Smaller values decrease memory
usage but increase runtime. (default: '5000')
(min: '1')
-PeptideQuantification:extract:mz_window <value> M/z window size for chromatogram extraction (uni
t: ppm if 1 or greater, else Da/Th) (default:
'10.0') (min: '0.0')
-PeptideQuantification:extract:IM_window <value> Ion mobility (IM) window for chromatogram extrac
tion in the IM dimension. Set to 0.0 to disable
IM filtering (even if data contains IM informati
on). The window is applied as +/- IM_window/2
around the median IM value of identified peptide
s. This parameter is automatically ignored if
the input data does not contain IM information
(determined via IMTypes::determineIMFormat).
...
for quality control. (default: '0.06') (min:
'0.0')
Parameters for detecting features in extracted ion chromatograms:
-PeptideQuantification:detect:mapping_tolerance <value> RT tolerance (plus/minus) for mapping peptide
IDs to features. Absolute value in seconds if 1
or greater, else relative to the RT span of the
feature. (default: '0.0') (min: '0.0')
Parameters for fitting exp. mod. Gaussians to mass traces.:
-PeptideQuantification:EMGScoring:max_iteration <number> Maximum number of iterations for EMG fitting.
(default: '100') (min: '1')
-PeptideQuantification:EMGScoring:init_mom <choice> Alternative initial parameters for fitting throu
gh method of moments. (default: 'true') (valid:
'true', 'false')
Parameters for FAIMS data processing:
-PeptideQuantification:faims:merge_features <choice> For FAIMS data with multiple compensation voltag
es: Merge features that represent the same analy
te detected at different CVs. Features are merge
d if they have the same charge and are within 5
seconds RT and 0.05 Da m/z. Intensities are summ
ed. (default: 'true') (valid: 'true', 'false')
Alignment:
-Alignment:model_type <choice> Options to control the modeling of retention
time transformations from data (default: 'b_spli
ne') (valid: 'linear', 'b_spline', 'lowess',
'interpolated')
Alignment:model:
-Alignment:model:type <choice> Type of model (default: 'b_spline') (valid: 'lin
ear', 'b_spline', 'lowess', 'interpolated')
Parameters for 'linear' model:
-Alignment:model:linear:symmetric_regression Perform linear regression on 'y - x' vs. 'y +
x', instead of on 'y' vs. 'x'.
-Alignment:model:linear:x_weight <choice> Weight x values (default: 'x') (valid: '1/x',
'1/x2', 'ln(x)', 'x')
-Alignment:model:linear:y_weight <choice> Weight y values (default: 'y') (valid: '1/y',
'1/y2', 'ln(y)', 'y')
-Alignment:model:linear:x_datum_min <value> Minimum x value (default: '1.0e-15')
-Alignment:model:linear:x_datum_max <value> Maximum x value (default: '1.0e15')
-Alignment:model:linear:y_datum_min <value> Minimum y value (default: '1.0e-15')
-Alignment:model:linear:y_datum_max <value> Maximum y value (default: '1.0e15')
Parameters for 'b_spline' model:
-Alignment:model:b_spline:wavelength <value> Determines the amount of smoothing by setting
the number of nodes for the B-spline. The number
is chosen so that the spline approximates a
low-pass filter with this cutoff wavelength.
The wavelength is given in the same units as
the data; a higher value means more smoothing.
'0' sets the number of nodes to twice the number
of input points. (default: '0.0') (min: '0.0')
-Alignment:model:b_spline:num_nodes <number> Number of nodes for B-spline fitting. Overrides
'wavelength' if set (to two or greater). A lower
value means more smoothing. (default: '5') (min
: '0')
-Alignment:model:b_spline:extrapolate <choice> Method to use for extrapolation beyond the origi
nal data range. 'linear': Linear extrapolation
using the slope of the B-spline at the correspon
ding endpoint. 'b_spline': Use the B-spline (as
for interpolation). 'constant': Use the constant
value of the B-spline at the corresponding endp
oint. 'global_linear': Use a linear fit through
the data (which will most probably introduce
discontinuities at the ends of the data range).
(default: 'linear') (valid: 'linear', 'b_spline'
, 'constant', 'global_linear')
-Alignment:model:b_spline:boundary_condition <number> Boundary condition at B-spline endpoints: 0 (val
ue zero), 1 (first derivative zero) or 2 (second
derivative zero) (default: '2') (min: '0' max:
'2')
Parameters for 'lowess' model:
-Alignment:model:lowess:span <value> Fraction of datapoints (f) to use for each local
regression (determines the amount of smoothing)
. Choosing this parameter in the range .2 to .8
usually results in a good fit. (default: '0.6666
66666666667') (min: '0.0' max: '1.0')
-Alignment:model:lowess:auto_span If true, or if 'span' is 0, automatically select
LOWESS span by cross-validation.
-Alignment:model:lowess:auto_span_min <value> Lower bound for auto-selected span. (default:
'0.15') (min: '1.0e-03')
-Alignment:model:lowess:auto_span_max <value> Upper bound for auto-selected span. (default:
'0.8') (max: '0.99')
-Alignment:model:lowess:auto_min_neighbors <number> Minimum number of neighbors (span*n) enforced
in auto mode. (default: '5') (min: '3')
-Alignment:model:lowess:auto_k_folds <number> K-folds for CV when n>50 (else LOO is used).
(default: '5') (min: '2')
-Alignment:model:lowess:auto_metric <choice> Metric for CV selection: one of {'p90','p95','p9
9','rmse','mae'}. (default: 'mae') (valid: 'p90'
, 'p95', 'p99', 'rmse', 'mae')
-Alignment:model:lowess:auto_span_grid <text> Optional explicit grid of span candidates in
(0,1]. Comma-separated list, e.g. '0.2,0.3,0.5'.
If empty, a default grid is used.
-Alignment:model:lowess:num_iterations <number> Number of robustifying iterations for lowess
fitting. (default: '3') (min: '0')
-Alignment:model:lowess:delta <value> Nonnegative parameter which may be used to save
computations (recommended value is 0.01 of the
range of the input, e.g. for data ranging from
1000 seconds to 2000 seconds, it could be set
to 10). Setting a negative value will automatica
lly do this. (default: '-1.0')
-Alignment:model:lowess:interpolation_type <choice> Method to use for interpolation between datapoin
ts computed by lowess. 'linear': Linear interpol
ation. 'cspline': Use the cubic spline for inter
polation. 'akima': Use an akima spline for inter
polation (default: 'cspline') (valid: 'linear',
'cspline', 'akima')
-Alignment:model:lowess:extrapolation_type <choice> Method to use for extrapolation outside the data
range. 'two-point-linear': Uses a line through
the first and last point to extrapolate. 'four-p
oint-linear': Uses a line through the first and
second point to extrapolate in front and and a
line through the last and second-to-last point
in the end. 'global-linear': Uses a linear regre
ssion to fit a line through all data points and
use it for interpolation. (default: 'four-point-
linear') (valid: 'two-point-linear', 'four-point
-linear', 'global-linear')
Parameters for 'interpolated' model:
-Alignment:model:interpolated:interpolation_type <choice> Type of interpolation to apply. (default: 'cspli
ne') (valid: 'linear', 'cspline', 'akima')
-Alignment:model:interpolated:extrapolation_type <choice> Type of extrapolation to apply: two-point-linear
: use the first and last data point to build a
single linear model, four-point-linear: build
two linear models on both ends using the first
two / last two points, global-linear: use all
points to build a single linear model. Note that
global-linear may not be continuous at the bord
er. (default: 'two-point-linear') (valid: 'two-p
oint-linear', 'four-point-linear', 'global-linea
r')
Alignment:align_algorithm:
-Alignment:align_algorithm:score_type <text> Name of the score type to use for ranking and
filtering (.oms input only). If left empty, a
score type is picked automatically.
-Alignment:align_algorithm:min_run_occur <number> Minimum number of runs (incl. reference, if any)
in which a peptide must occur to be used for
the alignment.
Unless you have very few runs or identifications
, increase this value to focus on more informati
ve peptides. (default: '2') (min: '2')
-Alignment:align_algorithm:max_rt_shift <value> Maximum realistic RT difference for a peptide
(median per run vs. reference). Peptides with
higher shifts (outliers) are not used to compute
the alignment.
If 0, no limit (disable filter); if > 1, the
final value in seconds; if <= 1, taken as a frac
tion of the range of the reference RT scale.
(default: '0.1') (min: '0.0')
-Alignment:align_algorithm:use_adducts <choice> If IDs contain adducts, treat differently adduct
ed variants of the same molecule as different.
(default: 'true') (valid: 'true', 'false')
Linking:
-Linking:nr_partitions <number> How many partitions in m/z space should be used
for the algorithm (more partitions means faster
runtime and more memory efficient execution).
(default: '100') (min: '1')
-Linking:min_nr_diffs_per_bin <number> If IDs are used: How many differences from match
ing IDs should be used to calculate a linking
tolerance for unIDed features in an RT region.
RT regions will be extended until that number
is reached. (default: '50') (min: '5')
-Linking:min_IDscore_forTolCalc <value> If IDs are used: What is the minimum score of
an ID to assume a reliable match for tolerance
calculation. Check your current score type! (def
ault: '1.0')
-Linking:noID_penalty <value> If IDs are used: For the normalized distances,
how high should the penalty for missing IDs be?
0 = no bias, 1 = IDs inside the max tolerances
always preferred (even if much further away).
(default: '0.0') (min: '0.0' max: '1.0')
Distance component based on m/z differences:
-Linking:distance_MZ:max_difference <value> Never pair features with larger m/z distance
(unit defined by 'unit') (default: '10.0') (min:
'0.0')
-Linking:distance_MZ:unit <choice> Unit of the 'max_difference' parameter (default:
'ppm') (valid: 'Da', 'ppm')
ProteinQuantification:
-ProteinQuantification:method <choice> - top - quantify based on three most abundant
peptides (number can be changed in 'top').
- iBAQ (intensity based absolute quantification)
, calculate the sum of all peptide peak intensit
ies divided by the number of theoretically obser
vable tryptic peptides (https://rdcu.be/cND1J).
Warning: only consensusXML or featureXML input
is allowed! (default: 'top') (valid: 'top', 'iBA
Q')
-ProteinQuantification:best_charge Distinguish between fraction and charge states
in detailed peptide output. For protein quantifi
cation, select one charge per modified peptide
globally: maximize the number of (fraction group
, label) assays with a positive abundance, then
break ties by total abundance; retain that charg
e's values in every assay.
By default, protein abundances are summed over
...
'fractions:aggregate', not by this flag.
Additional options for custom quantification using top N peptides.:
-ProteinQuantification:top:N <number> Calculate protein abundance from this number of
proteotypic peptides (most abundant first; '0'
for all) (default: '3') (min: '0')
-ProteinQuantification:top:aggregate <choice> Aggregation method used to compute protein abund
ances from peptide abundances (default: 'median'
) (valid: 'median', 'mean', 'weighted_mean',
'sum')
Options for combining the fractions of a fraction group.:
-ProteinQuantification:fractions:aggregate <choice> How the fractions of one fraction group are comb
ined into that group's (fraction group, label)
assay values.
- sum - add up every fraction, i.e. treat them
as the parts of one separated sample that they
are.
- best - keep a single fraction per peptide and
fraction group and discard the others. The fract
...
on and always report every file. (default: 'sum'
) (valid: 'sum', 'best')
Additional options for consensus maps (and identification results comprising multiple runs):
-ProteinQuantification:consensus:normalize Scale peptide abundances so that the median of
each (fraction group, label) assay matches the
overall median.
Abundances of zero count as 'not detected' and
are left out of the medians; an assay without
any positive abundance takes no part in the norm
alization.
-ProteinQuantification:consensus:fix_peptides Use the same peptides for protein quantification
across all (fraction group, label) assays.
With 'N 0',all peptides that occur in every assa
y are considered.
Otherwise ('N'), the N peptides that occur in
the most assays (independently of each other)
are selected,
breaking ties by total abundance (there is no
...
), not a measurement of absence.
PipEcho:
-PipEcho:fdr <value> MBR FDR threshold (0.05=5%). (default: '0.05')
(min: '0.0' max: '1.0')
-PipEcho:random_seed <number> Seed for the random number generator used to
select decoy donors. A fixed seed makes results
reproducible. (default: '0')
PipEcho:distance_MZ:
-PipEcho:distance_MZ:max_difference <value> Never pair features with larger m/z distance
(unit defined by 'unit') (default: '10.0') (min:
'0.0')
-PipEcho:distance_MZ:unit <choice> Unit of the 'max_difference' parameter (default:
'ppm') (valid: 'Da', 'ppm')
Common TOPP options:
-ini <file> Use the given TOPP INI file
-threads <n> Sets the number of threads allowed to be used
by the TOPP tool (0 = all available cores) (defa
ult: '1')
-write_ini <file> Writes the default configuration file
--help Shows options
--helphelp Shows all options (including advanced)
INI file documentation of this tool:
This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.
Each parameter entry contains the following information:
Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.
Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.