![]() |
OpenMS
|
Compute peptide and protein abundances from annotated feature/consensus maps or from identification results.
| potential predecessor tools | → ProteinQuantifier → | potential successor tools |
|---|---|---|
| IDMapper | external tools e.g. for statistical analysis | |
| FeatureLinkerUnlabeled (or another feature grouping tool) |
Reference:
Weisser et al.: An automated pipeline for high-throughput label-free quantitative proteomics (J. Proteome Res., 2013, PMID: 23391308).
Input: featureXML or consensusXML
Quantification is based on the intensity values of the features in the input files. Feature intensities are first accumulated to peptide abundances, according to the peptide identifications annotated to the features/feature groups. Then, abundances of the peptides of a protein are aggregated to compute the protein abundance.
The peptide-to-protein step uses the (e.g. 3) most abundant proteotypic peptides per protein to compute the protein abundances. This is a general version of the "top 3 approach" (but only for relative quantification) described in:
Silva et al.: Absolute quantification of proteins by LCMSE: a virtue of parallel MS acquisition (Mol. Cell. Proteomics, 2006, PMID: 16219938).
Only features/feature groups with unambiguous peptide annotation are used for peptide quantification. It is possible to resolve ambiguities before applying ProteinQuantifier using one of several equivalent mechanisms in OpenMS: IDConflictResolver, ConsensusID (algorithm best), or FileFilter (option id:keep_best_score_id).
Similarly, only proteotypic peptides (i.e. those matching to exactly one protein) are used for protein quantification by default. Peptide/protein IDs from multiple identification runs can be handled, but will not be differentiated (i.e. protein accessions for a peptide will be accumulated over all identification runs). See section "Optional input: Protein inference/grouping results" below for exceptions to this.
Peptides with the same sequence, but with different modifications are quantified separately on the peptide level, but treated as one peptide for the protein quantification (i.e. the contributions of differently-modified variants of the same peptide are accumulated).
Output granularity: assays vs. files and channels
By default one protein and peptide abundance is reported per assay. An assay is the experimental-design pair (fraction_group, label): it spans every fraction file of that fraction group at that label, and its reported value aggregates over those files. Columns are named abundance_fgroupF_labelL, where F is the design's Fraction_Group and L its Label. The SampleSection remains metadata and may group several assays as technical or biological replicates; ProteinQuantifier does not sum those replicates.
How the fractions of a group are combined into its assay values is controlled by fractions:aggregate. The default sum adds them up, treating them as the parts of one separated sample that they are. best instead keeps one fraction per peptide and fraction group and discards the rest: the fraction with the most labels at a positive abundance wins, ties are broken by the total of those abundances and then by the lower fraction number. The choice is made once for the whole fraction group and never per label - taking one channel from one fraction and another channel from a different fraction would mix physical aliquots and destroy the reporter-ion ratios that isobaric quantification consists of. Note that best reports a fraction of the material rather than all of it, which matters for LFQ, where the value is an absolute intensity, more than for isobaric data, where quantification is relative within a run and any single fraction preserves the ratios.
With file_and_channel_level_output the protein abundances are instead reported per (file, channel) cell. These cells are computed with the same peptide-level policy as the assay values (all peptidoforms are accumulated into one peptide; all charge states contribute by default, or only each peptidoform's selected charge with best_charge), but the peptide selection and the aggregation are applied per file. Two consequences are worth knowing:
top:N 0 together with top:aggregate sum. Top-N selection, median, mean and weighted_mean do not commute with aggregation across fractions, so for those settings the cells of an assay neither sum nor average to the assay value.top:N requirement ("at least N peptides") is likewise enforced per file, not per assay. In a fractionated experiment this is considerably stricter than the assay-level rule: a protein can easily have N peptides in an assay while no individual fraction contains N of them, in which case the protein is quantified at the assay level but all of its (file, channel) cells are reported as 0. Use top:N 0 (optionally with top:aggregate sum) or top:include_all if per-file values are wanted for such data.fractions:aggregate does not apply to them. A (file, channel) cell is one fraction by definition, so every file is always reported, even under best where the assay value comes from a single fraction. The two granularities then describe the data at different completeness on purpose.With best_charge, one charge is selected globally for each modified peptide. Charges are ranked first by the number of distinct assays with a positive abundance and then, on a tie, by total abundance across all assays (an exact tie deterministically keeps the lower charge). Every observation of the selected charge is retained and then combined over the fractions of an assay according to fractions:aggregate; the detailed peptide output still reports all observed fraction and charge combinations. The same selected-charge policy is used for assay and file/channel protein quantities.
The detailed peptide_out table that this flag produces has one row per (fraction, charge) and one abundance column per (file, channel) covering every file of the experimental design, so a row reports 0.0 for the files and channels its fraction does not cover. Without the flag, peptide_out instead writes one column per assay and one row per peptide, with fraction reported as "all".
Input: idXML
Quantification based on identification results uses spectral counting, i.e. the abundance of each peptide is the number of times that peptide was identified from an MS2 spectrum (considering only the best hit per spectrum). Different identification runs in the input become distinct inferred assays; this makes it possible to quantify several related runs at once by merging the corresponding idXML files with IDMerger. Depending on the presence of multiple runs, output format and applicable parameters are the same as for featureXML and consensusXML, respectively.
The notes above regarding quantification on the protein level and the treatment of modifications also apply to idXML input. In particular, this means that the settings top 0 and aggregate sum should be used to get the "classical" spectral counting quantification on the protein level (where all identifications of all peptides of a protein are summed up).
Optional input: Protein inference/grouping results
By default only proteotypic peptides (i.e. those matching to exactly one protein) are used for protein quantification. However, this limitation can be overcome: Protein inference results for the complete data set can be supplied with the protein_groups option (or included in a featureXML input). In that case, the peptide-to-protein references from that file are used (rather than those from in), and groups of indistinguishable proteins will be quantified. Each reported protein quantity then refers to the total for the respective group.
In order for everything to work correctly, it is important that the protein inference results come from the same identifications that were used to annotate the quantitative data. We suggest to use the OpenMS tool ProteinInference ProteinInference.
More information below the parameter specification.
Optional output: QPX Parquet (out_qpx)
out_qpx writes the quantification as a QPX collection - quantms.feature.parquet, quantms.psm.parquet and quantms.pg.parquet - for consensusXML input. QPX is an interchange format with a strict value contract, and OpenMS refuses to write a table it cannot represent rather than emit one that will not join. A refusal aborts the tool and leaves no files behind, including any view already written.
Consensus maps produced by ProteomicsLFQ and IsobaricWorkflow satisfy the contract by construction; those two are the supported producers. A map assembled by a different pipeline may not. What the contract requires, and how to satisfy it, is documented in one place: on OpenMS::QPXValueValidation, the class that enforces it.
The command line parameters of this tool are:
ProteinQuantifier -- Compute peptide and protein abundances
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ProteinQuantifier.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
+ Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.
Usage:
ProteinQuantifier <options>
Options (mandatory options marked with '*'):
-in <file>* Input file (valid formats: 'featureXML', 'consensusXML', 'idXML')
-protein_groups <file> Protein inference results for the identification runs that were
used to annotate the input (e.g. via the ProteinInference tool).
Information about indistinguishable proteins will be used for prot
ein quantification. (valid formats: 'idXML')
-design <file> Input file containing the experimental design (valid formats: 'tsv
')
-out <file> Output file for protein abundances (valid formats: 'csv')
-peptide_out <file> Output file for peptide abundances (valid formats: 'csv')
-mztab <file> Output file (mzTab) (valid formats: 'mzTab')
-out_qpx <directory> Output directory for QPX Parquet files (quantms.feature.parquet,
quantms.psm.parquet, quantms.pg.parquet). Only supported for conse
nsusXML input.
QPX has a strict value contract; input that does not meet it is
refused outright and no files are written. Maps produced by Proteo
micsLFQ or IsobaricWorkflow satisfy it by construction, other pipe
lines may not. The contract is documented on the OpenMS::QPXValueV
alidation class, which enforces it.
-method <choice> - top - quantify based on three most abundant peptides (number
can be changed in 'top').
- iBAQ (intensity based absolute quantification), calculate the
sum of all peptide peak intensities divided by the number of theor
etically observable tryptic peptides (https://rdcu.be/cND1J). Warn
ing: only consensusXML or featureXML input is allowed! (default:
'top') (valid: 'top', 'iBAQ')
-best_charge Distinguish between fraction and charge states in detailed peptide
output. For protein quantification, select one charge per modifie
d peptide globally: maximize the number of (fraction group, label)
assays with a positive abundance, then break ties by total abunda
nce; retain that charge's values in every assay.
By default, protein abundances are summed over all charge states.
How the retained values of several fractions are combined is gover
ned by 'fractions:aggregate', not by this flag.
Additional options for custom quantification using top N peptides.:
-top:N <number> Calculate protein abundance from this number of proteotypic peptid
es (most abundant first; '0' for all) (default: '3') (min: '0')
-top:aggregate <choice> Aggregation method used to compute protein abundances from peptide
abundances (default: 'median') (valid: 'median', 'mean', 'weighte
d_mean', 'sum')
-top:include_all Include results for proteins with fewer proteotypic peptides than
indicated by 'N' (no effect if 'N' is 0 or 1)
Options for combining the fractions of a fraction group.:
-fractions:aggregate <choice> How the fractions of one fraction group are combined into that
group's (fraction group, label) assay values.
- sum - add up every fraction, i.e. treat them as the parts of
one separated sample that they are.
- best - keep a single fraction per peptide and fraction group
and discard the others. The fraction is chosen ONCE per peptide,
ranked by the number of labels in which it has a positive abundanc
e and then by the total of those abundances (an exact tie keeps
...
are per fraction by definition and always report every file. (def
ault: 'sum') (valid: 'sum', 'best')
Additional options for consensus maps (and identification results comprising multiple runs):
-consensus:normalize Scale peptide abundances so that the median of each (fraction grou
p, label) assay matches the overall median.
Abundances of zero count as 'not detected' and are left out of
the medians; an assay without any positive abundance takes no part
in the normalization.
-consensus:fix_peptides Use the same peptides for protein quantification across all (fract
ion group, label) assays.
With 'N 0',all peptides that occur in every assay are considered.
Otherwise ('N'), the N peptides that occur in the most assays (ind
ependently of each other) are selected,
breaking ties by total abundance (there is no guarantee that the
best co-ocurring peptides are chosen!).
A peptide counts as occurring in an assay only where its abundance
...
measurement of absence.
-greedy_group_resolution <choice> Pre-process identifications with greedy resolution of shared pepti
des based on the protein group probabilities. (Only works with an
idXML file given as protein_groups parameter). (default: 'false')
(valid: 'true', 'false')
-file_and_channel_level_output <choice> Output protein abundances with detailed file+channel level headers
(similar to detailed peptide output). When enabled, protein outpu
t will show abundance_filename_channel columns instead of assay
columns.
Note that peptide selection and aggregation are then applied per
file, not per assay: 'top:N' requires N peptides in that single
file (much stricter than the assay-level rule for fractionated
data, where all cells of a quantified protein can end up 0), and
the cells only decompose the assay-level value for 'top:N' 0 with
'top:aggregate' sum. (default: 'false') (valid: 'true', 'false')
Output formatting options:
-format:separator <sep> Character(s) used to separate fields; by default, the 'tab' charac
ter is used
-format:quoting <method> Method for quoting of strings: 'none' for no quoting, 'double'
for quoting with doubling of embedded quotes,
'escape' for quoting with backslash-escaping of embedded quotes
(default: 'double') (valid: 'none', 'double', 'escape')
-format:replacement <x> If 'quoting' is 'none', used to replace occurrences of the separat
or in strings before writing (default: '_')
Common TOPP options:
-ini <file> Use the given TOPP INI file
-threads <n> Sets the number of threads allowed to be used by the TOPP tool (0
= all available cores) (default: '1')
-write_ini <file> Writes the default configuration file
--help Shows options
--helphelp Shows all options (including advanced)
INI file documentation of this tool:
This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.
Each parameter entry contains the following information:
Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.
Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.
Output format
The output files produced by this tool have a table format, with columns as described below:
Protein output (one protein/set of indistinguishable proteins per line):
top).(F, L). There is one self-describing column per assay in the experimental design.Peptide output (one peptide or - if best_charge is set - one charge state and fraction of a peptide per line):
best_charge was set.(F, L). If the charge in the preceding column is 0, this is the total abundance over all charge states; otherwise, it is only the abundance observed for the indicated charge (in this case, the detailed table uses file/channel columns instead). For consensusXML input, the reported values are already normalized if consensus:normalize was set.Protein quantification examples
While quantification on the peptide level is fairly straight-forward, a number of options influence quantification on the protein level - especially for consensusXML input. The three parameters top:N, top:include_all and consensus:fix_peptides determine which peptides are used to quantify proteins in different assays.
As an example, consider a protein with four proteotypic peptides. Each peptide is detected in a subset of three assays, as indicated in the table below. The peptides are ranked by abundance (1: highest, 4: lowest; assuming for simplicity that the order is the same in all assays).
| assay 1 | assay 2 | assay 3 | |
| peptide 1 | X | X | |
| peptide 2 | X | X | |
| peptide 3 | X | X | X |
| peptide 4 | X | X |
Different parameter combinations lead to different quantification scenarios, as shown here:
| parameters "*": no effect in this case | peptides used for quantification "(...)": not quantified here because ... | explanation | ||||
top | include_all | c.:fix_peptides | assay 1 | assay 2 | assay 3 | |
| 0 | * | no | 1, 2, 3, 4 | 2, 3, 4 | 1, 3 | all peptides |
| 1 | * | no | 1 | 2 | 1 | single most abundant peptide |
| 2 | * | no | 1, 2 | 2, 3 | 1, 3 | two most abundant peptides |
| 3 | no | no | 1, 2, 3 | 2, 3, 4 | (too few peptides) | three most abundant peptides |
| 3 | yes | no | 1, 2, 3 | 2, 3, 4 | 1, 3 | three or fewer most abundant peptides |
| 4 | no | * | 1, 2, 3, 4 | (too few peptides) | (too few peptides) | four most abundant peptides |
| 4 | yes | * | 1, 2, 3, 4 | 2, 3, 4 | 1, 3 | four or fewer most abundant peptides |
| 0 | * | yes | 3 | 3 | 3 | all peptides present in every assay |
| 1 | * | yes | 3 | 3 | 3 | single peptide present in most assays |
| 2 | no | yes | 1, 3 | (peptide 1 missing) | 1, 3 | two peptides present in most assays |
| 2 | yes | yes | 1, 3 | 3 | 1, 3 | two or fewer peptides present in most assays |
| 3 | no | yes | 1, 2, 3 | (peptide 1 missing) | (peptide 2 missing) | three peptides present in most assays |
| 3 | yes | yes | 1, 2, 3 | 2, 3 | 1, 3 | three or fewer peptides present in most assays |
Further considerations for parameter selection
With best_charge and the protein aggregation settings, there is a trade-off between comparability of protein abundances within an assay and of abundances for the same protein across different assays.
Setting best_charge may increase reproducibility between assays, but will distort the proportions of protein abundances within an assay. The reason is that ionization properties vary between peptides, but should remain constant across assays. Filtering by charge state can help to reduce the impact of feature detection differences between assays.
For aggregate, there is a qualitative difference between (intensity weighted) mean/median and sum in the effect that missing peptide abundances have (only if include_all is set or top is 0): (intensity weighted) mean and median ignore missing cases, averaging only present values. If low-abundant peptides are not detected in some assays, the computed protein abundances for those assays may thus be too optimistic. sum implicitly treats missing values as zero, so this problem does not occur and comparability across assays is ensured. However, with sum the total number of peptides ("summands") available for a protein may affect the abundances computed for it (depending on top), so results within an assay may become unproportional.