![]() |
OpenMS
|
Complete quantification workflow for MS1-labeled (SILAC, Dimethyl, ...) LC-MS/MS experiments.
| pot. predecessor tools | → MS1LabeledWorkflow → | pot. successor tools |
|---|---|---|
| PercolatorAdapter | MzTabExporter | |
| IDFilter |
This tool runs the complete quantification of an experiment whose samples were labeled before the LC-MS measurement so that the light and heavy forms of a peptide appear as separate MS1 features with a fixed mass shift: SILAC (Lys4/Lys6/Lys8, Arg6/Arg10), Dimethyl and ICPL labeling, in duplex, triplex or higher plex. It is the MS1-labeling counterpart of ProteomicsLFQ (label-free) and IsobaricWorkflow (TMT / iTRAQ reporter ions) and combines the standalone tools FeatureFinderMultiplex, IDMapper, IDConflictResolver, MultiplexResolver, MapAlignerIdentification, FeatureLinkerUnlabeledQT, ProteinInference and ProteinQuantifier into one run.
Input
in), with unique basenames. Native RAW input requires a build with WITH_THERMO_RAW and a .NET 8+ runtime. MS1 peaks are used for quantification; spectrum metadata are also read to recover identification FAIMS CVs. Profile and centroided data are both accepted (see algorithm:spectrum_type).ids), one file per spectra file in the same order, already filtered at PSM level (e.g. q-value < 0.01) and carrying Posterior Error Probability scores, e.g. produced with PercolatorAdapter (-score_type pep) or IDPosteriorErrorProbability, followed by IDFilter. Results from several search engines must be combined with ConsensusID rather than concatenated.labels), in the syntax of FeatureFinderMultiplex, e.g. [][Lys8,Arg10] for SILAC, [][Lys4,Arg6][Lys8,Arg10] for triple SILAC, [Dimethyl0][Dimethyl6] for Dimethyl. Every bracket is one channel; the channels are numbered from 1 in this order and are the Label column of the experimental design. Channels must be in increasing mass order, and SILAC requires an explicit unlabelled first channel [].design) with the columns Fraction_Group, Fraction, Spectra_Filepath, Label and Sample (see ExperimentalDesign). One row per (file, channel). Without a design every file is an unfractionated fraction group and every (file, channel) is its own sample.fasta). The identifications are then re-indexed with PeptideIndexer, which annotates protein sequences (for coverage), the decoy status and the theoretical peptide uniqueness needed by -protein_quantification strictly_unique_peptides.FAIMS multiplets, identification mapping, blacklist checks and linking are restricted to the same compensation voltage. CVs share the physical run's RT alignment and sample columns; they contribute separate multiplet evidence to peptide quantification. Identification CVs are recovered from the original spectrum references (including the preceding acquisition CV when MS2 metadata omit it). In a multi-CV run, an ID without a resolvable spectrum reference must carry a valid FAIMS_CV annotation; it cannot be assigned using RT alone. A single-CV run allows an unambiguous fallback to that voltage. Missing references are repaired only against MS2 spectra at the known CV, within 0.01 seconds of the ID's RT.
Important: the labels have to be part of the database search as (variable) modifications, e.g. Label:13C(6)15N(2) (K) and Label:13C(6)15N(4) (R) for Lys8/Arg10. Otherwise the MS2 spectra of the labeled channels stay unidentified and the observed mass shifts cannot be reconciled with the peptide sequences, so most multiplets end up as conflicts. The tool checks the search parameters recorded in ids for the modifications implied by labels and refuses to run if they are missing (use -force to proceed anyway).
Workflow
algorithm and label_mass_shifts), mapping of the identifications onto the multiplets (id_mapping), reduction to one identification per multiplet, and consolidation of quantitative and sequence information (OpenMS::MultiplexResolverAlgorithm, resolver): multiplets whose mass shifts contradict the labels found in the annotated sequence are removed from quantification (their identifications are kept for protein inference), incomplete multiplets are completed with dummy features (intensity 0 = absent, NaN = not quantifiable). Multiplets without identification are dropped, unless match_between_runs is set (see below). Once the resolver has used the labels, every identification is reduced to the peptide identity: the label modifications of labels are removed from the sequence, because the label belongs to the channel, not to the peptide, and the light and the heavy spectrum of one peptide have to name one peptide for linking, match between runs, inference and quantification (the convention MaxQuant uses as well). The label state stays documented on every identification as the meta values MS1Label:labeled_sequence (the peptidoform as searched), MS1Label:removed_labels (e.g. Lys8, or none) and MS1Label:channel (the 1-based channel the spectrum belongs to, i.e. the Label of the experimental design). The values also sit on every quantified consensus feature, for the identification it is quantified under. mzTab reports them as opt_global_* columns of the peptide and PSM sections, the QPX feature and psm views as cv_params. PSM-level output describes the spectrum match and therefore reports the peptidoform as searched (mzTab PSM section, QPX psm view), feature-level output the peptide identity (see OpenMS::MS1LabelState). The column headers describe every channel's labels in channel_description. A spectrum match that was mapped onto several multiplets stays on the one whose matched channel is closest to the precursor; distinct spectra of one peptide on distinct multiplets are all kept, and their channel values add up per peptide and charge in the quantification.alignment, identification-based, aligned to the run that shares the most identifications with every other run) and linking of the multiplets across runs (linking); the channels of every run are kept as sub-features, so the linked map has one column per (run, channel). Fractions are linked separately and then combined column-wise, exactly like ProteomicsLFQ does; a fraction measured in a single run is passed through. With match_between_runs, unidentified multiplets take part in the linking and take over the identification of a multiplet at the same position in another run (the SILAC equivalent of ProteomicsLFQ's -targeted_only false); multiplets that stay unidentified are not quantified.protein_inference), protein (and optionally PSM/peptide) FDR filtering, and peptide and protein quantification (ProteinQuantification), where the fractions of a fraction group are aggregated according to the design.Ratios. A labeled experiment measures its channels in one run, so its quantity is their ratio, and the tool computes it the way MaxQuant does, as a median of ratios rather than as a ratio of aggregated intensities (ratios, computed by MS1LabeledRatioQuantifier next to this tool):
ratios:reference_channel, the light one by default). Only positive, finite channels take part: an absent (dummy, intensity 0) or not-quantifiable channel is no measurement of a ratio.ratios:min_ratio_count peptides upwards (MaxQuant's "min. ratio count", 2 by default), with the number of contributing peptides next to it. Every ratio is also reported divided by the median peptide ratio of its (fraction group, channel), i.e. normalized on the assumption that most peptides do not change.The reference channel is reported with the ratio 1.0 it has by construction, wherever another channel was measured against it, so that every annotation covers the complete set of channels.
The ratios are annotated on the consensus features (MS1Label:evidence_ratio*, MS1Label:peptide_ratio*) and on the protein groups, so they reach the consensusXML and the mzTab peptide section (as opt_global_* columns). In the QPX pg view, whose rows are one per (protein group, fraction group, channel), they are written as that row's additional_intensities, named ratio and ratio_normalized under the row's own channel label; the number of contributing peptides sits in cv_params as ratio_count, being a count rather than an intensity. No separate ProteinQuantifier run is needed for any of it.
Next to the ratios, the per-channel abundances are reported as before (mzTab peptide and protein sections, QPX intensities): the summed peptide intensities per channel, like MaxQuant's Intensity columns. Dividing two of those is a ratio of aggregates, a different statistic from the ratios above, which weights peptides by their intensity. ProteinQuantification:consensus:normalize scales every assay to the overall median, which for a labeled experiment forces the median channel ratio to 1; leave it off unless that is intended.
max_nr_labelled_aas is used for both the feature detection and the resolver: it is the maximum number of labelled amino acids per peptide minus one, i.e. for tryptic SILAC the number of allowed missed cleavages. It should agree with the missed-cleavage setting of the search.
Output (at least one required; each output is optional individually)
out_cxml)out)out_qpx): quantms.feature.parquet, quantms.psm.parquet, quantms.pg.parquet. The channels are reported with the canonical SDRF/QPX labels (SILAC light, SILAC medium, SILAC heavy, DIMETHYL0, ...). Labels outside this vocabulary (ICPL, Leu3, plain mass shifts) cannot be exported to QPX; the tool refuses out_qpx for them up front.The command line parameters of this tool are:
MS1LabeledWorkflow -- Quantification workflow for MS1-labeled (SILAC, Dimethyl, ...) LC-MS/MS experiments.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_MS1LabeledWorkflow.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
+ Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.
Usage:
MS1LabeledWorkflow <options>
Options (mandatory options marked with '*'):
-in <file list>* Input: spectra files (mzML, or Thermo RAW with WITH_TH
ERMO_RAW and .NET 8+), one per LC-MS run. Only MS1
peaks are used; profile and centroided data are accept
ed. (valid formats: 'mzML', 'raw')
-ids <file list>* Identifications filtered at PSM level (e.g., q-value
< 0.01), one per spectra file in the same order.
The identifications must carry Posterior Error Probabi
lity scores (e.g. PercolatorAdapter with -score_type
pep,
or IDPosteriorErrorProbability) and the labels must
have been searched as (variable) modifications.
Combine results from several search engines with Conse
nsusID rather than concatenating them. (valid formats:
'idXML', 'mzId', 'idparquet')
-design <file> Experimental design (Fraction_Group, Fraction, Spectra
_Filepath, Label, Sample), one row per (file, channel)
.
'Label' is the 1-based position of the channel in '-la
bels'. If not given, every file is an unfractionated
fraction group and every (file, channel) is a separate
sample. (valid formats: 'tsv')
-fasta <file> Protein database. If given, the identifications are
re-indexed (PeptideIndexer): protein sequences (for
coverage),
decoy annotation and theoretical peptide uniqueness
(needed by 'strictly_unique_peptides') are taken from
it. (valid formats: 'fasta', 'fa', 'faa')
-labels <text> Labels used for labelling the samples, one bracket
per channel. [...] specifies the labels for a single
sample. For example
[][Lys8,Arg10] ... SILAC
[][Lys4,Arg6][Lys8,Arg10] ... triple-SILAC
[Dimethyl0][Dimethyl6] ... Dimethyl
[Dimethyl0][Dimethyl4][Dimethyl8] ... triple
...
' column of the experimental design). (default: '[][Ly
s8,Arg10]')
-max_nr_labelled_aas <int> Maximum number of labelled amino acids per peptide,
minus one. Peptides with up to (this value + 1) labell
ed amino acids
are considered by feature detection and resolver. For
SILAC with trypsin digestion, this is the maximum numb
er of missed cleavages. (default: '0') (min: '0')
-out <file> Optional output mzTab file. At least one output must
be specified. (valid formats: 'mzTab')
-out_cxml <file> Optional output consensusXML file. At least one output
must be specified. (valid formats: 'consensusXML')
-out_qpx <directory> Optional output directory for QPX Parquet files (quant
ms.feature.parquet, quantms.psm.parquet, quantms.pg.pa
rquet). At least one output must be specified.
-proteinFDR <threshold> Protein FDR threshold (0.05=5%). (default: '0.05')
(min: '0.0' max: '1.0')
-picked_proteinFDR <choice> Use a picked protein FDR? (default: 'false') (valid:
'true', 'false')
-psmFDR <threshold> FDR threshold for sub-protein level (e.g. 0.05=5%).
Use -FDR_type to choose the level. Cutoff is applied
at the highest level. (default: '1.0') (min: '0.0'
max: '1.0')
-FDR_type <option> Sub-protein FDR level. PSM, PSM+peptide (best PSM q-va
lue). (default: 'PSM') (valid: 'PSM', 'PSM+peptide')
-protein_inference <option> Infer proteins:
aggregation = aggregates all peptide scores across a
protein (using the best score)
bayesian = computes a posterior probability for
every protein based on a Bayesian network. (default:
'aggregation') (valid: 'aggregation', 'bayesian')
-match_between_runs <option> True: keep multiplets without an identification, so
that linking can hand them the identification of a
multiplet at the
same position in another run (the counterpart of Prote
omicsLFQ's '-targeted_only false').
false: only identified multiplets are quantified.
Cannot be combined with 'algorithm:knock_out': the
channel order of an unidentified multiplet is only
known from its detection pattern. (default: 'false')
(valid: 'true', 'false')
Parameters of the multiplet detection (FeatureFinderMultiplex):
-algorithm:charge <text> Range of charge states in the sample, i.e. min charge
: max charge. (default: '1:4')
-algorithm:rt_typical <value> Typical retention time [s] over which a characteristic
peptide elutes. (This is not an upper bound. Peptides
that elute for longer will be reported.) (default:
'40.0') (min: '0.0')
-algorithm:rt_band <value> The algorithm searches for characteristic isotopic
peak patterns, spectrum by spectrum. For some low-inte
nsity peptides, an important peak might be missing in
one spectrum but be present in one of the neighbouring
ones. The algorithm takes a bundle of neighbouring
spectra with width rt_band into account. For example
with rt_band = 0, all characteristic isotopic peaks
have to be present in one and the same spectrum. As
rt_band increases, the sensitivity of the algorithm
but also the likelihood of false detections increases.
(default: '0.0') (min: '0.0')
-algorithm:rt_min <value> Lower bound for the retention time [s]. (Any peptides
seen for a shorter time period are not reported.) (def
ault: '2.0') (min: '0.0')
-algorithm:mz_tolerance <value> M/z tolerance for search of peak patterns. (default:
'6.0') (min: '0.0')
-algorithm:mz_unit <choice> Unit of the 'mz_tolerance' parameter. (default: 'ppm')
(valid: 'Da', 'ppm')
-algorithm:intensity_cutoff <value> Lower bound for the intensity of isotopic peaks. (defa
ult: '1000.0') (min: '0.0')
-algorithm:peptide_similarity <value> Two peptides in a multiplet are expected to have the
same isotopic pattern. This parameter is a lower bound
on their similarity. (default: '0.5') (min: '-1.0'
max: '1.0')
-algorithm:averagine_similarity <value> The isotopic pattern of a peptide should resemble the
averagine model at this m/z position. This parameter
is a lower bound on similarity between measured isotop
ic pattern and the averagine model. (default: '0.4')
(min: '-1.0' max: '1.0')
Parameters for mapping the identifications onto the multiplets (IDMapper):
-id_mapping:rt_tolerance <value> RT tolerance (in seconds) for the matching (default:
'5.0') (min: '0.0')
-id_mapping:mz_tolerance <value> M/z tolerance (in ppm or Da) for the matching (default
: '20.0') (min: '0.0')
-id_mapping:mz_measure <choice> Unit of 'mz_tolerance' (ppm or Da) (default: 'ppm')
(valid: 'ppm', 'Da')
Parameters of the identification-based retention time alignment (MapAlignerIdentification):
-alignment:min_run_occur <number> Minimum number of runs (incl. reference, if any) in
which a peptide must occur to be used for the alignmen
t.
Unless you have very few runs or identifications, incr
ease this value to focus on more informative peptides.
(default: '2') (min: '2')
-alignment:max_rt_shift <value> Maximum realistic RT difference for a peptide (median
per run vs. reference). Peptides with higher shifts
(outliers) are not used to compute the alignment.
If 0, no limit (disable filter); if > 1, the final
value in seconds; if <= 1, taken as a fraction of the
range of the reference RT scale. (default: '0.1') (min
: '0.0')
Parameters for linking the multiplets across runs (FeatureLinkerUnlabeledQT):
-linking:nr_partitions <number> How many partitions in m/z space should be used for
the algorithm (more partitions means faster runtime
and more memory efficient execution). (default: '100')
(min: '1')
-linking:min_nr_diffs_per_bin <number> If IDs are used: How many differences from matching
IDs should be used to calculate a linking tolerance
for unIDed features in an RT region. RT regions will
be extended until that number is reached. (default:
'50') (min: '5')
-linking:min_IDscore_forTolCalc <value> If IDs are used: What is the minimum score of an ID
to assume a reliable match for tolerance calculation.
Check your current score type! (default: '1.0')
-linking:noID_penalty <value> If IDs are used: For the normalized distances, how
high should the penalty for missing IDs be? 0 = no
bias, 1 = IDs inside the max tolerances always preferr
ed (even if much further away). (default: '0.0') (min:
'0.0' max: '1.0')
Distance component based on m/z differences:
-linking:distance_MZ:max_difference <value> Never pair features with larger m/z distance (unit
defined by 'unit') (default: '10.0') (min: '0.0')
-linking:distance_MZ:unit <choice> Unit of the 'max_difference' parameter (default: 'ppm'
) (valid: 'Da', 'ppm')
Parameters of the peptide and protein abundances (ProteinQuantifier):
-ProteinQuantification:method <choice> - top - quantify based on three most abundant peptides
(number can be changed in 'top').
- iBAQ (intensity based absolute quantification), calc
ulate the sum of all peptide peak intensities divided
by the number of theoretically observable tryptic pept
ides (https://rdcu.be/cND1J). Warning: only consensusX
ML or featureXML input is allowed! (default: 'top')
(valid: 'top', 'iBAQ')
-ProteinQuantification:best_charge Distinguish between fraction and charge states in deta
iled peptide output. For protein quantification, selec
t one charge per modified peptide globally: maximize
the number of (fraction group, label) assays with a
positive abundance, then break ties by total abundance
; retain that charge's values in every assay.
By default, protein abundances are summed over all
charge states. How the retained values of several frac
tions are combined is governed by 'fractions:aggregate
', not by this flag.
Additional options for custom quantification using top N peptides.:
-ProteinQuantification:top:N <number> Calculate protein abundance from this number of proteo
typic peptides (most abundant first; '0' for all) (def
ault: '0') (min: '0')
-ProteinQuantification:top:aggregate <choice> Aggregation method used to compute protein abundances
from peptide abundances (default: 'sum') (valid: 'medi
an', 'mean', 'weighted_mean', 'sum')
Options for combining the fractions of a fraction group.:
-ProteinQuantification:fractions:aggregate <choice> How the fractions of one fraction group are combined
into that group's (fraction group, label) assay values
.
- sum - add up every fraction, i.e. treat them as the
parts of one separated sample that they are.
- best - keep a single fraction per peptide and fracti
on group and discard the others. The fraction is chose
n ONCE per peptide, ranked by the number of labels in
...
report every file. (default: 'sum') (valid: 'sum',
'best')
Additional options for consensus maps (and identification results comprising multiple runs):
-ProteinQuantification:consensus:normalize Scale peptide abundances so that the median of each
(fraction group, label) assay matches the overall medi
an.
Abundances of zero count as 'not detected' and are
left out of the medians; an assay without any positive
abundance takes no part in the normalization.
-ProteinQuantification:consensus:fix_peptides Use the same peptides for protein quantification acros
s all (fraction group, label) assays.
With 'N 0',all peptides that occur in every assay are
considered.
Otherwise ('N'), the N peptides that occur in the most
assays (independently of each other) are selected,
breaking ties by total abundance (there is no guarante
e that the best co-ocurring peptides are chosen!).
...
ce.
Parameters of the channel ratios, the reported quantity of a labeled experiment:
-ratios:reference_channel <number> Channel the ratios are formed against, as the 'Label'
of the experimental design (1 = the light channel of
'-labels'). (default: '1') (min: '1')
-ratios:min_ratio_count <number> Minimum number of peptide ratios a protein group needs
before a ratio is reported for it (MaxQuant's 'min.
ratio count'). Groups below it are reported without a
ratio, not with a less certain one. (default: '2')
(min: '1')
-ratios:normalize <choice> Additionally report every ratio divided by the median
peptide ratio of its (fraction group, channel), i.e.
assuming that most peptides do not change. The unnorma
lized ratios are reported either way. (default: 'true'
) (valid: 'true', 'false')
Common TOPP options:
-ini <file> Use the given TOPP INI file
-threads <n> Sets the number of threads allowed to be used by the
TOPP tool (0 = all available cores) (default: '1')
-write_ini <file> Writes the default configuration file
--help Shows options
--helphelp Shows all options (including advanced)
INI file documentation of this tool:
This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.
Each parameter entry contains the following information:
Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.
Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.